Model optimizer, multi-hop question and answer model training method and device, and multi-hop question and answer method and device

By combining Hamiltonian mechanics model optimizer, dynamically adjusting the step size and maintaining the geometric structure of the Hamiltonian system, the shortcomings of the existing multi-hop question-and-answer system in the processing of complex inference chains are solved, and the stability and efficiency of model training are significantly improved.

CN119940554AActive Publication Date: 2025-05-06BEIJING SANSAN SMART EDUCATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510110558.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing multi-hop question and answer system shows difficulty in dealing with complex inference chains, and although the integration of knowledge graphs enhances structured reasoning capabilities, it sacrifices a certain degree of flexibility and makes it difficult to move freely in complex knowledge structures.

Method used

A model optimizer is proposed, combining Hamiltonian mechanics, through the model inference framework, Hamiltonian equation module, Cyclone integrator and function optimization module, the step length is dynamically adjusted to maintain the geometric structure integrity of the Hamiltonian system, and optimize model parameters to improve the stability and efficiency of the inference path.

Benefits of technology

It significantly improves the stability and efficiency of model training, can maintain the geometric integrity of the inference path in complex knowledge structures, and improves the performance and interpretability of multi-hop question and answer systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940554A_ABST
    Figure CN119940554A_ABST
Patent Text Reader

Abstract

The invention provides a model optimizer, a multi-hop question and answer model training method and device, a multi-hop question and answer method and device, electronic equipment and a computer readable storage medium, and relates to the technical field of data processing, in particular to the technical fields of deep learning, natural language processing and the like. According to the specific implementation scheme, the model optimizer comprises a model reasoning framework used for obtaining model parameters of a model and calculating the gradient of a loss function about the model parameters; the Hamiltonian equation module is used for obtaining momentum based on the gradient and the Hamiltonian equation; the symplectic integrator is used for obtaining updated model parameters through a symplectic integral algorithm based on the momentum and maintaining a geometric structure of a reasoning path of the model; and the function optimization module is used for minimizing the curvature and the torsion degree of the reasoning path.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the field of data processing technology, and particularly relates to technical fields such as deep learning and natural language processing, and in particular to a model optimizer, a multi-hop question-answering model training method and device, a multi-hop question-answering method and device, an electronic device, and a computer-readable storage medium. Background Art

[0002] Multi-hop question answering (QA) is a natural language processing task in which answering a question requires reasoning and synthesis from multiple information sources or multiple facts. Unlike traditional single-hop QA (which only requires extracting the answer from one information source), multi-hop QA requires the model to be able to connect different pieces of information and perform multi-step reasoning to arrive at the final answer. In multi-hop QA, users usually ask a complex question, and the model needs to not only identify relevant information, but also understand the relationship between this information and perform reasoning.

[0003] Multi-hop question answering is challenging because it requires deep semantic understanding and reasoning, often involving logic, causality, and comprehensive processing of information. Multi-hop question answering systems require models to be able to navigate complex knowledge structures and to cleverly connect scattered pieces of information to form a logically rigorous and coherent answer. This ability is crucial for developing AI systems that can reason in simple terms, provide easy-to-understand results, and cope with problems that often do not have direct, one-step solutions in the real world. Summary of the invention

[0004] The present disclosure provides a model optimizer, a multi-hop question-answering model training method and device, a multi-hop question-answering method and device, an electronic device, and a computer-readable storage medium.

[0005] According to a first aspect, a model optimizer is provided, which includes: a model reasoning framework, used to obtain model parameters of a model and calculate the gradient of a loss function with respect to the model parameters; a Hamiltonian equation module, used to obtain momentum based on the gradient and the Hamiltonian equation, wherein the momentum is used to characterize the change in reasoning between continuous model parameters; a symplectic integrator, used to obtain updated model parameters based on the momentum through a symplectic integration algorithm, and maintain the geometric structure of the reasoning path when the model parameters are obtained to update the model parameters; and a function optimization module, used to minimize the curvature and tortuosity of the reasoning path.

[0006] According to the second aspect, a multi-hop question-answering model training method is provided, the method comprising: obtaining a training data set, the training data set comprising at least one training data, the training data comprising: a question, an answer, and at least one related factual information; obtaining a multi-hop question-answering network; inputting the training data in the training data set into the multi-hop question-answering network to obtain a result output by the multi-hop question-answering network; training the multi-hop question-answering network based on a model optimizer and the result output by each multi-hop question-answering network to obtain a trained multi-hop question-answering network, wherein the model optimizer uses a model optimizer as described in any implementation method of the first aspect to adjust model parameters of the multi-hop question-answering network.

[0007] According to the third aspect, a multi-hop question and answer model training method is provided, the method comprising: obtaining questions to be answered and answer-related information, the answer-related information comprising: context information of the question to be answered and at least one of the knowledge graph information; a multi-hop question and answer model trained based on the questions to be answered, the answer-related information and the multi-hop question and answer model training method described in any implementation of the second aspect, to obtain the answer output by the multi-hop question and answer model.

[0008] According to a fourth aspect, a multi-hop question-answering model training device is provided, the device comprising: a sample acquisition unit, configured to acquire a training data set, the training data set comprising at least one training data, the training data comprising: a question, an answer and at least one related factual information; a network acquisition unit, configured to acquire a multi-hop question-answering network; a result acquisition unit, configured to input the training data in the training data set into the multi-hop question-answering network, and obtain a result output by the multi-hop question-answering network; an adjustment unit, configured to train the multi-hop question-answering network based on a model optimizer and the result output by each multi-hop question-answering network, and obtain a trained multi-hop question-answering network, wherein the model optimizer uses a model optimizer as described in any implementation method of the first aspect to adjust the model parameters of the multi-hop question-answering network.

[0009] According to the fifth aspect, a multi-hop question and answer device is provided, which includes: a question acquisition unit, configured to obtain questions to be answered and answer-related information, the answer-related information including: context information of the questions to be answered and at least one of the knowledge graph information; an answer acquisition unit, configured to obtain the answer output by the multi-hop question and answer model based on the questions to be answered, the answer-related information and the multi-hop question and answer model trained by the multi-hop question and answer model training device as described in any implementation method of the fourth aspect.

[0010] According to the sixth aspect, an electronic device is provided, which includes: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any implementation of the second aspect or the third aspect.

[0011] According to a seventh aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described in any implementation of the second aspect or the third aspect.

[0012] The model optimizer and model reasoning framework provided by the embodiments of the present disclosure are used to obtain the model parameters of the model and calculate the gradient of the loss function with respect to the model parameters; the Hamiltonian equation module is used to obtain momentum based on the gradient and the Hamiltonian equation, and the momentum is used to characterize the change of reasoning between continuous model parameters; the symplectic integrator is used to obtain the updated model parameters through the symplectic integration algorithm based on the momentum, and maintain the geometric structure of the reasoning path when the model parameters are updated; the function optimization module is used to minimize the curvature and torsion of the reasoning path. Thus, a model optimizer combined with Hamiltonian mechanics is provided, which can dynamically adjust the step size according to the Hamiltonian state, and the geometric structure integrity of the Hamiltonian system can be maintained through the symplectic integrator during the entire training process of the model, which significantly improves the stability and efficiency of model training.

[0013] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. Figure 1 is a structural schematic diagram of an embodiment of a model optimizer according to the present disclosure; Figure 2a is the phase diagram for the concentrated problem solving in the two-dimensional Hamiltonian system of this disclosure; Figure 2b is a phase diagram for multi-concept reasoning in a two-dimensional Hamiltonian system in the present disclosure; Figure 3 It is a structural diagram of the velocity, acceleration and angle changes of the trajectory described by the Vernet coordinate system; Figure 4 is a flowchart of an embodiment of a multi-hop question-answering model training method according to the present disclosure; Figure 5 is a flow chart of an embodiment of the multi-hop question-answering method according to the present disclosure; Figure 6 is a structural schematic diagram of an embodiment of a multi-hop question-answering model training device according to the present disclosure; Figure 7 is a structural schematic diagram of an embodiment of a multi-hop question-answering device according to the present disclosure; Figure 8It is a block diagram of an electronic device used to implement the model reasoning framework construction method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0015] Unless explicitly stated otherwise, throughout the specification and claims, the term “comprise” or variations such as “include” or “comprising”, etc., will be understood to include the stated elements or components but not to exclude other elements or components.

[0016] The technical solution of the present disclosure is described below by means of specific embodiments. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combination step, or other methods and steps may be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and are not used to limit the scope of the present disclosure. Unless otherwise specified, the numbering of each method step is only for the purpose of identifying each method step, and does not limit the order of arrangement of each method or limit the scope of implementation of the present disclosure. Changes or adjustments in their relative relationships can also be regarded as the scope of implementation of the present disclosure without substantial changes in the technical content.

[0017] The sources of the raw materials and instruments used in the examples are not particularly limited and can be purchased from the market or prepared according to conventional methods known to those skilled in the art.

[0018] Although current AI reasoning optimization technologies have made rapid progress, they still face many challenges. Transformer-based models, such as BERT (Bidirectional Encoder Representations from Transformers) and GPT (Generative Pre-trained Transformer), have revolutionized language understanding, but they often struggle to handle complex reasoning chains. The incorporation of knowledge graphs has enhanced the ability of structured reasoning, but at the expense of flexibility. Iterative optimization and advanced attention mechanisms have improved the accuracy of reasoning, but often at the expense of interpretability and computational efficiency. Reinforcement learning provides a promising path for reasoning optimization, but its effectiveness is overly dependent on the rationality of reward design. Neuro-symbolic methods attempt to combine the flexibility of neural networks with the clarity of symbolic reasoning, but face the problem of scalability. Recent innovations in prompt engineering and contextual learning have shown great potential, but their generalization ability still needs to be improved.

[0019] Although significant progress has been made in optimizing the inference process, many current methods are like black boxes, making it difficult to interpret their inference process or ensure their reliability. However, it is still necessary to build a more comprehensive theoretical framework to better integrate the understanding of different methods.

[0020] In response to the defects in traditional technologies, the present disclosure proposes a model optimizer, in which the reasoning process of artificial intelligence corresponds to the Hamiltonian system, thereby providing a new framework for in-depth analysis and optimization of complex artificial intelligence cognitive processes. Figure 1 A structural schematic diagram of an embodiment of a model optimizer according to the present disclosure is shown. The model optimizer 100 includes: a model reasoning framework 101, a Hamiltonian equation module 102, a symplectic integrator 103 and a function optimization module 104.

[0021] Among them, the model reasoning framework 101 is used to obtain the model parameters of the model and calculate the gradient of the loss function with respect to the model parameters. The Hamiltonian equation module 102 is used to obtain momentum based on the gradient and the Hamiltonian equation, and the momentum is used to characterize the change of reasoning between continuous model parameters. The symplectic integrator 103 is used to obtain the updated model parameters based on the momentum through the symplectic integration algorithm, and maintain the geometric structure of the reasoning path when the model parameters are obtained to update the model parameters. The function optimization module 104 is used to minimize the curvature and distortion of the reasoning path.

[0022] In this embodiment, the way in which the model inference framework 101 calculates the gradient through the model parameters and the loss function is a conventional method in the art and will not be repeated here.

[0023] Specifically, in Hamiltonian mechanics, the Hamiltonian H(q, p) combines the kinetic energy T(p) and the potential energy V(q), and the potential energy V(q) is related to the loss function, and the gradient calculated by the loss function is directly related to the potential energy.

[0024] In this embodiment, the model optimizer is used to optimize the training steps of the model. The update module parameters are calculated through the current model parameters of the model, and the update module parameters are output to the model to provide a model parameter update reference for the model. The model parameters obtained by the model reasoning framework 101 can be various types of models, such as the model is a multi-hop question-answering model.

[0025] In this embodiment, the reasoning path is a continuous model parameter change trajectory of the model, and the continuous model parameters are also model parameters to updated model parameters, wherein the updated model parameters are parameters after the model parameters are updated.

[0026] In this embodiment, the input of the model optimizer is the current model parameters of the model, and the output of the model optimizer is the updated model parameters. Its working principle is to calculate the Hamiltonian based on the momentum and gradient, and dynamically adjust the step size. The step size is an adaptive value calculated by the model optimizer based on the Hamiltonian energy state (kinetic energy and potential energy), which is used to control the amplitude of parameter update and ensure the symplectic structure (geometric characteristics) and energy conservation. It is an intermediate variable in the parameter update process; the momentum and parameters are updated through the symplectic integral rule to obtain the updated model parameters. The essence of updating the model parameters is to evolve the parameters from the current state to the next state through the symplectic integral rule, which strictly follows the Hamiltonian dynamics law.

[0027] In this embodiment, the model reasoning framework 101 and the Hamiltonian equation module 102 regard the reasoning process of the model as a trajectory in space, that is, a reasoning path, such as Figure 2a and Figure 2b As shown, in Figure 2a and Figure 2b The black solid line G represents the trajectory, S is the energy level of the inference, and Figure 2b , H is the key concept equilibrium point, and the movement of the model reasoning process follows the Hamiltonian equation, which is shown in Equations (1) and (2).

[0028] (1) (2) In formula (1) and formula (2), H R Denotes the Hamiltonian, and Equation (2) involves the gradient of the loss function with respect to the model parameter q, indicating that the gradient is indeed used in the update process. The model optimizer retains the geometric characteristics of the system, ensuring the stability and accuracy of the update process. It uses the gradient of the loss function to update the momentum p, which in turn affects the update of the model parameter q and obtains the updated model parameter. The gradient of the loss function is important for guiding the optimization process. They are integrated into the Hamiltonian dynamics to help update the momentum p and thus the model parameter q.

[0029] The symplectic integrator 103 is used to maintain the geometric structure of the inference path of the model reasoning framework when the model parameters are updated through the model parameters. Specifically, the symplectic integrator 103 ensures the geometric characteristics of the model parameter evolution (i.e., the inference path) through numerical methods. The function optimization module 104 is a component driving the optimization process independent of the symplectic integrator 103. The symplectic integrator 103 ensures that the inference path strictly follows the geometric laws of Hamiltonian dynamics (such as symplectic form conservation and energy approximation conservation), ensuring the long-term stability of the model inference path (such as the model energy error does not accumulate with the training steps). This is the basic framework of the optimizer and defines the legal boundary of the parameter update. Based on the legal trajectory defined by the symplectic integrator, the function optimization module 104 further optimizes the smoothness of the path by adding regularization terms (such as curvature and twisting), making the model's inference path more coherent without affecting the task performance. Its role is similar to "selecting a better route under the established physical laws." The function optimization module 104 constrains the geometric characteristics (such as curvature and twist) of the reasoning path through regularization terms, and its optimization goal is achieved on the legal parameter evolution manifold defined by the symplectic integrator, thereby improving the coherence and interpretability of the reasoning path without destroying the geometric structure of the system.

[0030] In this embodiment, the inference phase space of the model inference framework 101 and the Hamiltonian equation module 102 inherits the symplectic geometric structure, which perfectly maintains its geometric properties during the inference evolution process. This phenomenon is similar to the Liouville theorem in classical mechanics, which states that under the action of the Hamiltonian dynamic system, the volume of the phase space is constant: (3) In equation (3), ω represents the symplectic form, which is a basic concept in symplectic geometry and is used to describe the symmetry and structure in phase space. i is the position coordinate q i The differential of represents an infinitesimal change in position and direction in phase space. i is the momentum coordinate p i The differential of represents an infinitesimal change in the direction of momentum in phase space. represents the exterior product, which is an operation in multilinear algebra that defines the product of two differential forms. Equation (3) describes the definition of the symplectic form in the phase space, where the phase space is composed of position coordinates q i and momentum coordinate p i High-dimensional space.

[0031] In this embodiment, the symplectic integrator 103 approximately solves the Hamiltonian equation by numerical integration, ensuring that the evolution of the system state follows the law of Hamiltonian dynamics at each step of updating the model parameters, thereby maintaining the symplectic structure and energy conservation properties of the system. Specifically, the symplectic integrator 103 uses a custom symplectic integrator based on the Forest-Ruth algorithm, which is an efficient fourth-order symplectic integration algorithm.

[0032] In this embodiment, the symplectic integrator is an algorithm for numerically solving the Hamiltonian equation. Its core advantage is that it can maintain the symplectic structure of the system, thereby maintaining the stability of energy conservation and other important physical quantities during long-term integration. According to the Hamiltonian equation, the update rule of the model optimizer is expressed as formula (4): (4) In formula (4), represents a step size that is adaptively adjusted according to the current Hamiltonian value. Through this Hamiltonian perspective, the advanced analysis framework in physics is introduced into artificial intelligence reasoning, similar to the application of the Vernet coordinate system. With the help of this model reasoning framework, the Hamiltonian equation module and the new integrator, the function optimization module 104 determines the curvature of the reasoning path. and distortion : (5) In formula (5), is the trajectory curve, the first-order derivative γ'(t) represents the tangent vector of the path, that is, the inference direction and speed; the second-order derivative γ''(t) represents the normal acceleration of the path; in formula (5), × represents the vector cross product. Figure 3 The trajectory curve shown in the figure is used to describe the velocity, acceleration and angle change of the trajectory when the trajectory curve moves. Figure 3 middle, , They are tangent and normal respectively. Specifically, is the tangent line along the curve, is the normal line pointing to the center of the curve. The trajectory curve can also include: and The perpendicular binormal unit vector.

[0033] In this embodiment, the function optimization module 104 calculates the curvature and torsion of the reasoning path to quantify the geometric characteristics of the reasoning path. These geometric characteristics can be used to guide the step size adjustment of the reasoning path, so that the symplectic integrator 103 can generate updated model parameters. The curvature and torsion of the reasoning path are geometric characteristics that describe the movement of the reasoning process in the phase space. The curvature describes the degree to which the curve deviates from the straight line, that is, it reflects the degree of curvature of the reasoning path; while the torsion (also called torsion) describes the degree to which the curve deviates from the plane. In three-dimensional space, torsion is an important geometric quantity of the curve, which reflects the degree of "twisting" of the curve in space. The present disclosure reflects the degree of twisting of the reasoning path.

[0034] In this embodiment, in high-dimensional space, the concept in three-dimensional space can be used to understand the "distortion" of the reasoning path. In other words, the reasoning path is regarded as a curve γ(t) in high-dimensional space, and its distortion can be calculated by the following steps Step_1 to Step_4: Step 1. Assume that the reasoning path γ(t) is a function of time t, where t can represent the step or time step of reasoning.

[0035] Step 2. Calculate the first and second order derivatives.

[0036] Step 3. Calculate the curvature based on formula (5).

[0037] Step_4. Calculate the torsion according to the torsion formula.

[0038] In this embodiment, high curvature may indicate that the reasoning path has undergone drastic changes in certain steps, which may correspond to important turning points or key reasoning steps in the reasoning process. The geometric characteristics of the reasoning path (such as curvature) may reflect changes in cognitive flexibility in the reasoning process. Curvature is a geometric quantity that describes the degree of curvature of a curve at a certain point. The greater the curvature, the more drastic the curvature of the curve at that point. High torsion may indicate that the reasoning path has undergone complex distortions in high-dimensional space, which may correspond to the integration of information in multiple dimensions or complex logical relationships involved in the reasoning process.

[0039] In this embodiment, by introducing the concepts of curvature and torsion, a new geometric perspective is provided for analyzing and optimizing the reasoning path, which helps to more deeply understand the reasoning mechanism of the model and further improve the performance of the model.

[0040] These geometric properties reveal the "cognitive flexibility" of the system, where high curvature may indicate a rapid change in reasoning paths. At the same time, conservation laws in the reasoning process can also be explored. For example, quantities similar to angular momentum in physical systems may symbolize the constant elements in the effective reasoning process: (6) In classical mechanics, L represents the angular momentum of the system. In equation (6), q is the position vector and p is the momentum vector. Angular momentum is a conserved quantity that represents the inertia of the system in rotational motion. In Hamiltonian mechanics, the conservation of angular momentum is determined by the symmetry of the system. Specifically, it is caused by the symmetry of the system under rotation.

[0041] In this embodiment, the concept of angular momentum is introduced into the reasoning process of artificial intelligence. Here, q can be understood as the embedding vector of the reasoning state (i.e., the model parameter), and p is the momentum vector, which represents the change of the reasoning state. Therefore, the above formula (6) can be interpreted as the "cognitive angular momentum" in the reasoning process, which reflects a certain form of conservation or invariance in the reasoning process.

[0042] In this embodiment, the model optimizer is a symplectic optimizer, which aims to optimize the energy consumption and path efficiency during the reasoning process. The goals of the symplectic optimizer are: 1. In the Hamiltonian dynamic framework (the above-mentioned model reasoning framework and Hamiltonian equation module), energy consists of kinetic energy and potential energy. The step size of the reasoning path is adjusted by the model optimizer to update the model parameters so that the total energy consumption is minimized, thereby improving the reasoning efficiency.

[0043] 2. In classical mechanics, angular momentum is a conserved quantity. In the Hamiltonian dynamic framework, by maintaining the conservation of angular momentum, the stability of certain key information or reasoning paths in the reasoning process can be ensured, avoiding unnecessary fluctuations or errors in the reasoning process.

[0044] 3. By optimizing the step size of the reasoning path, the reasoning process is made more coherent and explainable, thereby improving the transparency of the model and user trust.

[0045] By placing AI reasoning in the Hamiltonian framework, an innovative path for in-depth analysis and optimization is opened up, which can quantify the "energy" consumption of cognitive processes, explore the geometric structure of reasoning paths, and potentially reveal the basic laws that govern efficient cognition. This method not only provides an advanced optimization technique, but also provides a new perspective for understanding and improving the reasoning ability of AI, which is particularly critical for dealing with complex tasks such as multi-hop question answering.

[0046] In this embodiment, the model reasoning framework 101 of the model optimizer can calculate the gradient of the loss function with respect to the model parameters in each iterative training of the model. The gradient indicates the direction in which the loss function changes fastest, and the opposite direction is the direction in which the model parameters are updated to ensure that the value of the loss function decreases.

[0047] In this embodiment, based on the calculated gradient, the model optimizer will update the parameters of the model according to a certain step size (learning rate). This step size determines the amplitude of the parameter update. If it is too small, it may lead to slow convergence, while if it is too large, it may cause oscillation or even divergence.

[0048] In this embodiment, after obtaining the model parameters of the current iterative training of the model, the model optimizer obtains updated model parameters under the joint action of the model reasoning framework, Hamiltonian equation model, symplectic integrator and function optimization module, and outputs the updated model parameters to the model so that the model continues to be trained in the next iterative training until the model meets the complete training conditions.

[0049] In this embodiment, the symplectic integrator 103 dynamically adjusts the step size according to the current Hamiltonian state and the curvature and tortuosity of the reasoning path to obtain updated model parameters. Specifically, when the curvature of the reasoning path is large, it may indicate that the reasoning process is undergoing a rapid change phase. At this time, the optimizer can appropriately reduce the step size to ensure the stability of the reasoning; on the contrary, when the curvature is small, the step size can be appropriately increased to speed up the reasoning. By adjusting the step size, the model optimizer aims to minimize the loss function of the reasoning trajectory, which may include the accuracy of the reasoning, the smoothness of the path, and other related factors.

[0050] In this embodiment, in addition to using the symplectic integrator, the model optimizer also considers the adaptive adjustment of the step length of the model reasoning path, and the calculation of the step length is based on the value of the current Hamiltonian: size= \frac{\text{lr}}{\sqrt{\text{hamiltonian}} + \epsilon} (7) In formula (7), lr: learning rate, which is a hyperparameter used to control the update speed of model parameters in deep learning or machine learning. Hamiltonian: Hamiltonian, usually used to describe the total energy of a physical system, but in the field of optimization and machine learning, this term may be used to represent a measure related to the objective function or loss function, which may be a measure of the complexity of the problem. Epsilon: a small constant, usually used to avoid the denominator being zero and ensure the stability of numerical calculations.

[0051] In equation (7), by adjusting the learning rate to adapt to the change of Hamiltonian, a more stable optimization process may be achieved. When the Hamiltonian is large, it means that the system may be in a more complex or high-energy state. At this time, the parameters can be updated more cautiously by reducing the step size (i.e., the adjustment value of the learning rate); conversely, when the Hamiltonian is small, the step size can be increased to accelerate convergence. In practical applications, such an adjustment method may need to be customized and debugged according to specific problems and models to achieve the best learning effect and convergence speed. Such a design also ensures that the integrity of the symplectic geometry structure is maintained during the training process.

[0052] In this embodiment, by modeling the reasoning process as a dynamic system in a high-dimensional phase space, the model can quickly find the optimal path in a complex knowledge graph. This optimization can effectively reduce unnecessary redundant steps in the reasoning process, thereby improving the response speed of the system.

[0053] The model optimizer provided in this embodiment dynamically adjusts the model parameters of the model through momentum and step size, so that it can improve the performance of complex reasoning tasks while following the Hamiltonian dynamics law (symplectic structure, energy conservation), and adjusts the model parameters through a physically inspired numerical method (symplectic integral) instead of adjusting the abstract "reasoning step size" or independently optimizing intermediate variables. It can improve the performance and interpretability of multi-hop question-answering tasks.

[0054] The model optimizer and model reasoning framework provided by the embodiments of the present disclosure are used to obtain the model parameters of the model and calculate the gradient of the loss function with respect to the model parameters; the Hamiltonian equation module is used to obtain momentum based on the gradient and the Hamiltonian equation, and the momentum is used to characterize the change of reasoning between continuous model parameters; the symplectic integrator is used to obtain the updated model parameters through the symplectic integration algorithm based on the momentum, and maintain the geometric structure of the reasoning path when the model parameters are updated; the function optimization module is used to minimize the curvature and torsion of the reasoning path. Thus, a model optimizer combined with Hamiltonian mechanics is provided, which can dynamically adjust the step size according to the Hamiltonian state, and the geometric structure integrity of the Hamiltonian system can be maintained through the symplectic integrator during the entire training process of the model, which significantly improves the stability and efficiency of model training.

[0055] In some embodiments of the present disclosure, the above-mentioned symplectic integrator adopts a Hamiltonian algorithm to obtain updated model parameters, and the Hamiltonian algorithm includes: the difference between a first sub-algorithm and a second sub-algorithm, wherein the first sub-algorithm is an algorithm for cognitive effort to change model parameters, the second sub-algorithm is an algorithm for correlation of model parameters, and momentum is related to the first sub-algorithm.

[0056] In this optional implementation, the Hamiltonian algorithm is used to characterize the Hamiltonian, which is the "energy" function in the reasoning process, including the "kinetic energy" and "potential energy" of the reasoning. The optimizer adjusts the momentum and position of the system to minimize the Hamiltonian, thereby optimizing the reasoning process.

[0057] In this optional implementation, the inference state (Model parameters) are represented in the form of vectors in a high-dimensional embedding space, as shown in Equation (8): (8) In formula (8), is an embedding function, and is the input text. Input text First, the natural language model's tokenizer is used to segment the words into a token. Each token are mapped to their corresponding embedding vectors ei: (9) As shown in formula (10), momentum Represents changes in reasoning between consecutive states: (10) Hamiltonian: (11) In formula (11), It is the “kinetic energy” or cognitive effort to change state , where μ is a fixed coefficient that can be set based on demand; and in formula (11), is the "potential energy" or relevance of the current state ,in is a similarity function (e.g., cosine similarity) and is the expected answer embed.

[0058] The model optimizer provided by this optional implementation adopts a Hamiltonian algorithm, which effectively characterizes the Hamiltonian, and the Hamiltonian effectively combines the model parameters and the Hamiltonian momentum, thereby improving the rationality of constructing the model reasoning framework.

[0059] Figure 4 A process 400 of an embodiment of a multi-hop question answering model training method according to the present disclosure is shown. The multi-hop question answering model training method includes the following steps: Step 401: Obtain a training data set.

[0060] In this embodiment, the execution subject on which the multi-hop question-answering model training method runs can obtain the training data set in a variety of ways. For example, the execution subject can obtain the training data set stored in the database server through a wired connection or a wireless connection. For another example, the user can obtain the training data set collected by the terminal by communicating with the terminal.

[0061] In this embodiment, the training data set includes at least one training data, and the training data includes: a question, an answer, and at least one relevant factual information. The at least one relevant factual information refers to at least two relevant factual texts used to reason and answer questions in a multi-hop question-answering task. The at least two factual texts are usually extracted from different information sources or knowledge bases and are used to build a reasoning chain from questions to answers. In multi-hop question-answering, the answer to a question often cannot be directly derived from a single fact, but requires reasoning based on multiple facts.

[0062] Specifically, at least two factual texts provide the necessary information basis for multi-hop reasoning. For example, suppose a question involves two related but not directly related facts - "Fact 1" and "Fact 2", "Fact 1" provides a relationship or condition, and "Fact 2" provides another relationship or condition. By combining these two facts, the model can reason to get the correct answer.

[0063] In the training data set, each training data includes: questions, answers, and at least one related factual information. These data are used to train the model to learn how to extract relevant information from multiple facts and make inferences. Through this training, the model can better simulate the human reasoning process when facing new multi-hop question-answering tasks, thereby improving the accuracy and coherence of the answers.

[0064] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of the training data set involved are carried out after authorization and comply with relevant laws and regulations.

[0065] Step 402: Obtain a multi-hop question-answering network.

[0066] In this embodiment, the multi-hop question-answering network is used to analyze the input question, determine the reasoning chain in multiple different facts, and obtain the answer. The multi-hop question-answering network can be a multimodal language model based on GPT-4o. GPT-4o is a multimodal model that can process multiple modal inputs including text (such as images, audio, etc.). In the disclosure, the ability of GPT-4o to process text information can be utilized, but its multimodal processing potential is retained.

[0067] In this embodiment, the multi-hop question-answering network is the initial network for analyzing the answer. Since the multi-hop question-answering network has not undergone the model process, the accuracy of the answer output by the multi-hop question-answering network cannot be guaranteed.

[0068] Step 403: input the training data in the training data set into the multi-hop question-answering network to obtain the output result of the multi-hop question-answering network.

[0069] In this embodiment, the execution subject can select training data from the training data set obtained in step 401, and execute the training steps from step 403 to step 404 to complete an iterative training of the multi-hop question-answering network. The selection method and the number of selected training data from the training data set are not limited in this application, and the number of iterative training of the multi-hop question-answering network is not limited. For example, in one iterative training, multiple continuous training data can be randomly selected, and the result output by the multi-hop question-answering network is obtained through the selected training data. Step 404: Based on the model optimizer and the results output by each multi-hop question-answering network, the multi-hop question-answering network is trained to obtain a trained multi-hop question-answering network.

[0070] In this embodiment, the model optimizer uses the model optimizer disclosed in the above embodiment to adjust the model parameters of the multi-hop question-answering network.

[0071] In this embodiment, during each iterative training of the multi-hop question-answering network, training data is selected from the training data set and input into the multi-hop question-answering network to obtain the answer output by the multi-hop question-answering network. In each iterative training, the model optimizer obtains the model parameters of the multi-hop question-answering network, and uses the model recommendation framework and symplectic integrator therein to update the model parameters to obtain updated model parameters, and applies the updated model parameters to the multi-hop question-answering network before the next iterative training.

[0072] In this embodiment, during each iterative training of the multi-hop question-answering network, the network loss value of the multi-hop question-answering network is calculated. Based on the network loss value and the number of iterative training, it is detected whether the multi-hop question-answering network meets the training completion conditions; if the training completion conditions are met, a trained multi-hop question-answering network is obtained.

[0073] In this embodiment, the loss function of the multi-hop question-answering network can adopt a cross-entropy loss function, which can measure the difference between two different probability distributions in the same random variable, and is represented as the difference between the true probability distribution and the predicted probability distribution in machine learning. The smaller the value of the cross-entropy loss function, the better the prediction effect of the multi-hop question-answering network.

[0074] The multi-hop question-answering model training method provided by the embodiments of the present disclosure first obtains a training data set, where the training data set includes at least one training data, and the training data includes: a question, an answer, and at least one related factual information; secondly, obtains a multi-hop question-answering network; then, inputs the training data in the training data set into the multi-hop question-answering network to obtain the output result of the multi-hop question-answering network; finally, based on the model optimizer and the output result of each multi-hop question-answering network, trains the multi-hop question-answering network to obtain the trained multi-hop question-answering network, adjusts the model parameters of the multi-hop question-answering network through the model optimizer, and improves the efficiency, stability, and interpretability of the decision-making process of artificial intelligence.

[0075] In some optional implementations of the present disclosure, the multi-hop question-answering network includes: a multimodal language subnet and a classification layer; the multimodal language subnet is used to characterize the correspondence between the multimodal material and the identified text results, and the classification layer is used to classify the text results into corresponding answers to obtain answer classification results. The inputting of the training data in the training data set into the multi-hop question-answering network to obtain the result output by the multi-hop question-answering network includes: selecting training data from the training data set; inputting the training data into the multi-hop question-answering network to obtain the answer classification result of the multi-hop question-answering network; the training of the multi-hop question-answering network based on the model optimizer and the result of each multi-hop question-answering network output to obtain the trained multi-hop question-answering network includes: calculating the network loss value of the multi-hop question-answering network based on the answer classification result and the predefined Hamiltonian loss function; training the multi-hop question-answering network based on the loss value and the model optimizer to obtain a trained multi-hop question-answering model.

[0076] In this optional implementation, the Hamiltonian loss function changes the reasoning path of the multi-hop question-answering network based on the change of the Hamiltonian in the Hamiltonian dynamics. Specifically, the Hamiltonian loss function can directly adopt the cross entropy function.

[0077] In this optional implementation, the above-mentioned calculation of the network loss value of the multi-hop question-answering network based on the answer classification result and the predefined Hamiltonian loss function includes: inputting the answer classification result into the Hamiltonian loss function to obtain the network loss value.

[0078] In this optional implementation, the multi-hop question-answering network is trained based on the loss value and the model optimizer to obtain a trained multi-hop question-answering model, including: based on the loss value, detecting whether the multi-hop question-answering network meets the training completion conditions; if the training completion conditions are not met, the model optimizer is used to update the model parameters of the multi-hop question-answering network, and the multi-hop question-answering network after the parameter update is continued to be detected based on the loss value to determine whether the multi-hop question-answering network meets the training completion conditions. Among them, the training completion conditions include: the training completion conditions include: the network loss value of the multi-hop question-answering network is less than the first loss value threshold. Among them, the first loss threshold can be determined based on specific training requirements, for example, the first loss threshold is 0.01.

[0079] In this optional implementation, the multimodal language subnet actually refers to a multimodal language model based on GPT-4o. GPT-4o is a multimodal model that can process multiple modal inputs including text (such as images, audio, etc.).

[0080] In this optional implementation, the classification layer is a custom layer attached to the multimodal language subnet, which is used to further classify the representations generated by the multimodal language subnet to adapt to specific multi-hop question answering tasks.

[0081] In this optional implementation, the multimodal language subnet is responsible for processing the input text information and generating corresponding semantic representations. These representations capture the rich semantic information of the input text and provide a basis for subsequent classification tasks. The classification layer is used to classify the semantic representations generated by the multimodal language subnet to determine the correct answer in the multi-hop question-answering task. Specifically, the output of the classification layer is the classification result for the multi-hop question-answering problem, such as determining whether a candidate answer is correct or selecting the most appropriate answer among multiple candidate answers.

[0082] In the present invention, the output of the classification layer is the classification result of the answer to the multi-hop question answering problem. Specifically, the classification layer may output a binary classification (e.g., correct or wrong), or select the most likely correct answer from multiple candidate answers. The specific design of the classification layer depends on the task requirements, but in the context of this disclosure, it is used to determine the answer to the multi-hop question answering problem.

[0083] The method for training a multi-hop question-answering network provided by this optional implementation manner calculates the network loss value based on the answer classification result output by the classification layer and the Hamiltonian loss function when the multi-hop question-answering network includes: a multimodal language subnetwork and a classification layer; based on the network loss value and a model optimizer, the multi-hop question-answering network is trained to obtain a trained multi-hop question-answering model, thereby improving the result output accuracy of the multi-hop question-answering model.

[0084] Optionally, the multi-hop question-answering network includes: a multimodal language subnet and a classification layer; the inputting of the training data in the training data set into the multi-hop question-answering network to obtain the result output by the multi-hop question-answering network includes: selecting training data from the training data set; inputting the training data into the multi-hop question-answering network to obtain the answer classification result of the multi-hop question-answering network; the training of the multi-hop question-answering network based on the model optimizer and the result output by each multi-hop question-answering network to obtain the trained multi-hop question-answering network includes: determining the current iteration training number of the multi-hop question-answering network; training the multi-hop question-answering network based on the iteration training number and the model optimizer to obtain the trained multi-hop question-answering model. Among them, the training of the multi-hop question-answering network based on the iteration training number and the model optimizer to obtain the trained multi-hop question-answering model includes: in response to detecting that the iteration training number is less than the training number threshold, obtaining updated model parameters based on the model optimizer, adjusting the multi-hop question-answering network using the updated model parameters, and continuing to train the multi-hop question-answering network.

[0085] In some optional implementations of the present disclosure, the multi-hop question-answering network is trained based on the loss value and the model optimizer to obtain a trained multi-hop question-answering model, including: based on the loss value, detecting whether the multi-hop question-answering network meets the training completion conditions; in response to detecting that the multi-hop question-answering network does not meet the training completion conditions, using the model optimizer to adjust the model parameters of the multi-hop question-answering network; continuing to select training data from the training data set, inputting the training data into the multi-hop question-answering network, and obtaining the answer classification results of the multi-hop question-answering network.

[0086] In this optional implementation, the training completion condition includes: the network loss value of the multi-hop question-answering network is less than a first loss value threshold, wherein the first loss threshold can be determined based on specific training requirements.

[0087] In this optional implementation, using a model optimizer to adjust the model parameters of the multi-hop question-answering network means: obtaining updated model parameters through the model optimizer, and replacing the current model parameters of the multi-hop question-answering network with the updated model parameters.

[0088] The method for training a multi-hop question-answering model provided by this optional implementation adopts a model optimizer to adjust the model parameters of the multi-hop question-answering network in response to detecting that the multi-hop question-answering network does not meet the training completion conditions; continues to select training data from the training data set, inputs the training data into the multi-hop question-answering network, and obtains the answer classification result of the multi-hop question-answering network, thereby providing a reliable implementation method for the training of the multi-hop question-answering network.

[0089] Optionally, the above-mentioned training of the multi-hop question-answering network based on the loss value and the model optimizer to obtain a trained multi-hop question-answering model also includes: in response to detecting that the multi-hop question-answering network meets the training completion condition, obtaining a trained multi-hop question-answering model.

[0090] In some optional implementations of the present disclosure, the above-mentioned Hamiltonian loss function includes: a cross-entropy loss function and a regularization term based on a parameter norm. Based on the answer classification result and the predefined Hamiltonian loss function, calculating the network loss value of the multi-hop question-answering network includes: calculating the cross-entropy loss value based on the answer classification result and the cross-entropy loss function; calculating the regularized loss value based on the answer classification result and the regularization term; and obtaining the network loss of the multi-hop question-answering network based on the cross-entropy loss value and the regularized loss value.

[0091] In this optional implementation, the Hamiltonian loss function includes not only the traditional classification loss, but also a regularization term based on the parameter norm to reduce the "energy" consumption in the Hamiltonian and prevent overfitting.

[0092] The Hamiltonian loss function provided by this optional implementation cleverly combines the traditional classification loss with the regularization term of the model parameter norm, which encourages the model to seek solutions that reduce "energy" consumption, thereby effectively promoting the model to develop more stable and generalizable representations. By penalizing such high parameter values, the risk of overfitting is reduced and the model's generalization ability in a variety of situations is improved.

[0093] In some optional implementations of the present disclosure, the network loss of the multi-hop question-answering network obtained based on the cross-entropy loss value and the regularization loss value includes: multiplying the cross-entropy loss value by the cross-entropy loss weight plus the regularization loss value by the regularization loss weight to obtain the network loss of the multi-hop question-answering network.

[0094] In this optional implementation, according to the proportion of the cross entropy function and the regularization term in model training, a cross entropy loss weight is set for the cross entropy function and a regularization loss weight is set for the regularization term. By setting the two weights, the reliability of the network loss of the multi-hop question-answering network is improved.

[0095] Optionally, obtaining the network loss of the multi-hop question-answering network based on the cross entropy loss value and the regularization loss value includes: adding the cross entropy loss value to the regularization loss value to obtain the network loss of the multi-hop question-answering network.

[0096] Figure 5 A process 500 of an embodiment of a multi-hop question-answering method according to the present disclosure is shown. The multi-hop question-answering method includes the following steps: Step 501: Obtain questions to be answered and answer-related information.

[0097] In this embodiment, the reply-related information includes: context information of the question to be replied and at least one of the knowledge graph information. The context information is information related to the question to be replied, and the knowledge graph information is professional information for replying the question to be replied.

[0098] In this embodiment, the multi-hop question-answering model processes the questions to be answered and the answer-related information, and the answer to the question to be answered can be obtained. The execution subject of the multi-hop question-answering method can obtain the questions to be answered and the answer-related information in a variety of ways. For example, the execution subject can obtain the questions to be answered and the answer-related information stored in the database server through a wired connection or a wireless connection. For another example, the execution subject can also receive the questions to be answered and the answer-related information collected in real time by the terminal or other devices.

[0099] In this embodiment, the answer to the question to be replied is a result of the question to be replied. Specifically, the result includes: selecting information from context information or knowledge graph information.

[0100] Step 502: Based on the question to be answered, the answer-related information and the multi-hop question-answering model, obtain the answer output by the multi-hop question-answering model.

[0101] In this embodiment, the multi-hop question-answering model is a model trained using the multi-hop question-answering model training method of the above embodiment.

[0102] In this embodiment, the execution entity can input the questions to be answered and the answer-related information obtained from step 501 into the multi-hop question-answering model, thereby obtaining the answers output by the multi-hop question-answering model.

[0103] In this embodiment, the multi-hop question answering model can be adopted as described above. Figure 4 The multi-hop question answering model trained by the method described in the embodiment is obtained. The specific training process can be found in Figure 4 The relevant description of the embodiments will not be repeated here.

[0104] The multi-hop question-answering model training method provided by the embodiments of the present disclosure first obtains questions to be answered and answer-related information, where the answer-related information includes: context information of the questions to be answered and at least one of the knowledge graph information; then, based on the questions to be answered, the answer-related information and the multi-hop question-answering model trained by the multi-hop question-answering model training method, the answer output by the multi-hop question-answering model is obtained, thereby improving the accuracy and reliability of the answers obtained.

[0105] Further references Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a multi-hop question-answering model training device. Figure 4 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0106] like Figure 6As shown, the multi-hop question-answering model training device 600 provided in this embodiment includes: a sample acquisition unit 601, a network acquisition unit 602, a result acquisition unit 603, and an adjustment unit 604. Among them, the above-mentioned sample acquisition unit 601 can be configured to acquire a training data set, and the training data set includes at least one training data, and the training data includes: a question, an answer, and at least one related factual information. The above-mentioned network acquisition unit 602 can be configured to acquire a multi-hop question-answering network. The above-mentioned result acquisition unit 603 can be configured to input the training data in the training data set into the multi-hop question-answering network to obtain the result output by the multi-hop question-answering network. The above-mentioned adjustment unit 604 can be configured to train the multi-hop question-answering network based on the model optimizer and the result output by each multi-hop question-answering network, and obtain the trained multi-hop question-answering network. The model optimizer uses the model optimizer provided by the above embodiment to adjust the model parameters of the multi-hop question-answering network.

[0107] In this embodiment, in the multi-hop question-answering model training device 600, the specific processing of the sample acquisition unit 601, the network acquisition unit 602, the result acquisition unit 603, and the adjustment unit 604 and the technical effects thereof can be referred to respectively. Figure 4 The relevant descriptions of step 401, step 402, step 403, and step 404 in the corresponding embodiment are not repeated here.

[0108] In some embodiments of the present disclosure, the multi-hop question-answering network unit is configured as: a multimodal language subnet and a classification layer; the multimodal language subnet is used to characterize the correspondence between the multimodal material and the identified text results, and the classification layer is used to classify the text results into corresponding answers to obtain answer classification results, and the result obtaining unit 603 is configured as: selecting training data from the training data set; inputting the training data into the multi-hop question-answering network to obtain the answer classification result of the multi-hop question-answering network; the adjustment unit 604 is configured as: based on the answer classification result and the predefined Hamiltonian loss function, calculating the network loss value of the multi-hop question-answering network; based on the loss value and the model optimizer, training the multi-hop question-answering network to obtain a trained multi-hop question-answering model.

[0109] In some embodiments of the present disclosure, the above-mentioned adjustment unit 604 is configured to: based on the loss value, detect whether the multi-hop question and answer network meets the training completion condition; in response to detecting that the multi-hop question and answer network does not meet the training completion condition, use a model optimizer to adjust the model parameters of the multi-hop question and answer network; continue to select training data from the training data set, input the training data into the multi-hop question and answer network, and obtain the answer classification result of the multi-hop question and answer network.

[0110] In some embodiments of the present disclosure, the above-mentioned Hamiltonian loss function includes: a cross entropy loss function and a regularization term based on a parameter norm, and the above-mentioned adjustment unit 604 is configured to: calculate the cross entropy loss value based on the answer classification result and the cross entropy loss function; calculate the regularized loss value based on the answer classification result and the regularization term; and obtain the network loss of the multi-hop question-answering network based on the cross entropy loss value and the regularized loss value.

[0111] In some embodiments of the present disclosure, the adjustment unit 604 is configured to: multiply the cross entropy loss value by the cross entropy loss weight and add the regularization loss value by the regularization loss weight to obtain the network loss of the multi-hop question-answering network.

[0112] The multi-hop question-answering model training device provided by the embodiment of the present disclosure, first, the sample acquisition unit 601 acquires a training data set, the training data set includes at least one training data, the training data includes: a question, an answer and at least one related factual information; secondly, the network acquisition unit 602 acquires a multi-hop question-answering network; then, the result acquisition unit 603 inputs the training data in the training data set into the multi-hop question-answering network to obtain the result output by the multi-hop question-answering network; finally, the adjustment unit 604 trains the multi-hop question-answering network based on the model optimizer and the result output by each multi-hop question-answering network to obtain the trained multi-hop question-answering network, thereby adjusting the model parameters of the multi-hop question-answering network through the model optimizer, thereby improving the efficiency, stability and interpretability of the decision-making process of artificial intelligence.

[0113] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a multi-hop question-answering device. Figure 5 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0114] like Figure 7 As shown, the multi-hop question-answering device 700 provided in this embodiment includes: a question acquisition unit 701, and an answer acquisition unit 702. Among them, the above-mentioned question acquisition unit 701 can be configured to obtain the question to be answered and the answer-related information, and the answer-related information includes: the context information of the question to be answered and at least one of the knowledge graph information. The above-mentioned answer acquisition unit 702 can be configured to obtain the answer output by the multi-hop question-answering model based on the question to be answered, the answer-related information and the multi-hop question-answering model trained by the multi-hop question-answering model training method.

[0115] In this embodiment, in the multi-hop question-answering model training device 700, the specific processing of the question obtaining unit 701 and the answer obtaining unit 702 and the technical effects thereof can be referred to in Figure 5 The relevant descriptions of step 501 and step 502 in the corresponding embodiment are not repeated here. According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0116] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their modes are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0117] like Figure 8 As shown, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0118] A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0119] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as a multi-hop question-answering model training method or a multi-hop question-answering method. For example, in some embodiments, the jump question-answering model training method or the multi-hop question-answering method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the jump question-answering model training method or the multi-hop question-answering method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the skip question answering model training method or the multi-hop question answering method in any other appropriate manner (for example, by means of firmware).

[0120] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable jump question-answering model training device or a multi-hop question-answering device, so that the program code, when executed by the processor or controller, implements the mode / operation specified in the flow chart and / or block diagram. The program code may be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.

[0122] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0124] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0125] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0126] The foregoing description of specific exemplary embodiments of the present disclosure is for the purpose of illustration and demonstration. These descriptions are not intended to limit the present disclosure to the precise form disclosed, and it is clear that many changes and variations can be made based on the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present disclosure and its practical application, so that those skilled in the art can realize and utilize various different exemplary embodiments of the present disclosure and various different selections and changes. The scope of the present disclosure is intended to be defined by the claims and their equivalents.

Claims

1. A model optimizer, the model optimizer comprising: A model inference framework, used to obtain model parameters of the model and calculate the gradient of the loss function with respect to the model parameters; A Hamiltonian equation module, used to obtain momentum based on the gradient and the Hamiltonian equation, wherein the momentum is used to characterize the change of reasoning between continuous model parameters; A symplectic integrator, configured to obtain updated model parameters based on the momentum by a symplectic integration algorithm, and to maintain a geometric structure of an inference path when the model parameters are used to obtain the updated model parameters; The function optimization module is used to minimize the curvature and distortion of the reasoning path.

2. The model optimizer according to claim 1, wherein the symplectic integrator adopts a Hamiltonian algorithm to obtain updated model parameters, and the Hamiltonian algorithm comprises: The difference between a first sub-algorithm and a second sub-algorithm, wherein the first sub-algorithm is an algorithm of cognitive effort to change model parameters and the second sub-algorithm is an algorithm of correlation of model parameters, the momentum being related to the first sub-algorithm.

3. A multi-hop question-answering model training method, the method comprising: Acquire a training data set, wherein the training data set includes at least one training data, and the training data includes: a question, an answer, and at least one related factual information; Get a multi-hop question answering network; Inputting the training data in the training data set into the multi-hop question answering network to obtain a result output by the multi-hop question answering network; Based on the model optimizer and the results output by the multi-hop question-answering network each time, the multi-hop question-answering network is trained to obtain a trained multi-hop question-answering network, and the model optimizer uses the model optimizer described in claim 1 or 2 to adjust the model parameters of the multi-hop question-answering network.

4. The method according to claim 3, wherein the multi-hop question-answering network comprises: Multimodal language subnet and classification layer; The multimodal language subnet is used to characterize the correspondence between the multimodal material and the recognized text results, the classification layer is used to classify the text results into corresponding answers to obtain answer classification results, and the training data in the training data set is input into the multi-hop question-answering network to obtain the result output by the multi-hop question-answering network, including: Selecting training data from the training data set; Inputting the training data into the multi-hop question-answering network to obtain an answer classification result of the multi-hop question-answering network; The training of the multi-hop question answering network based on the model optimizer and the result outputted by the multi-hop question answering network each time to obtain the trained multi-hop question answering network comprises: Based on the answer classification result and a predefined Hamiltonian loss function, calculating a network loss value of the multi-hop question-answering network; Based on the loss value and the model optimizer, the multi-hop question-answering network is trained to obtain a trained multi-hop question-answering model.

5. According to the method of claim 4, the step of training the multi-hop question-answering network based on the loss value and the model optimizer to obtain a trained multi-hop question-answering model comprises: Based on the loss value, detecting whether the multi-hop question-answering network meets a training completion condition; In response to detecting that the multi-hop question-answering network does not meet the training completion condition, using the model optimizer to adjust the model parameters of the multi-hop question-answering network; Continue to select training data from the training data set, input the training data into the multi-hop question-answering network, and obtain the answer classification result of the multi-hop question-answering network.

6. The method according to claim 4, wherein the Hamiltonian loss function comprises: A cross entropy loss function and a regularization term based on a parameter norm, wherein the network loss value of the multi-hop question-answering network is calculated based on the answer classification result and a predefined Hamiltonian loss function, including: Based on the answer classification result and the cross entropy loss function, calculating a cross entropy loss value; Calculate a regularization loss value based on the answer classification result and the regularization term; Based on the cross entropy loss value and the regularization loss value, a network loss of the multi-hop question-answering network is obtained.

7. According to the method of claim 6, obtaining the network loss of the multi-hop question-answering network based on the cross entropy loss value and the regularization loss value comprises: The network loss of the multi-hop question-answering network is obtained by multiplying the cross entropy loss value by the cross entropy loss weight and adding the regularization loss value by the regularization loss weight.

8. A multi-hop question answering method, the method comprising: Obtaining questions to be answered and answer-related information, wherein the answer-related information includes: at least one of context information of the question to be answered and knowledge graph information; Based on the question to be answered, the answer related information and the multi-hop question answering model trained by the multi-hop question answering model training method described in any one of claims 3-7, the answer output by the multi-hop question answering model is obtained.

9. A multi-hop question-answering model training device, the device comprising: A sample acquisition unit is configured to acquire a training data set, wherein the training data set includes at least one training data, and the training data includes: a question, an answer, and at least one related factual information; A network acquisition unit configured to acquire a multi-hop question-answering network; A result obtaining unit is configured to input the training data in the training data set into the multi-hop question answering network to obtain a result output by the multi-hop question answering network; The adjustment unit is configured to train the multi-hop question-answering network based on the model optimizer and the results output by the multi-hop question-answering network each time to obtain a trained multi-hop question-answering network, and the model optimizer uses the model optimizer described in claim 1 or 2 to adjust the model parameters of the multi-hop question-answering network.

10. A multi-hop question-answering device, comprising: A question acquisition unit is configured to acquire a question to be answered and answer related information, wherein the answer related information includes: at least one of context information of the question to be answered and knowledge graph information; The answer obtaining unit is configured to obtain the answer output by the multi-hop question and answer model based on the question to be answered, the answer related information and the multi-hop question and answer model trained by the multi-hop question and answer model training device according to claim 9.

11. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 3 to 8.

12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 3 to 8.

Citation Information

Patent Citations

  • Knowledge graph multi-hop question and answer method and model based on cognitive reasoning

    CN113360604A

  • Method and device for improving electromagnetic property measurement precision of equipment

    CN114611387A

  • Building health monitoring and evaluation method and system based on physical neural network

    CN119249073A

  • Generating and managing deep tensor neural networks

    US20200151580A1

Cited By

  • Rescue resource demand quantity prediction and dynamic supply method for major accident

    CN120235421A