Information processing method and electronic equipment

By obtaining target parameters and target models with deep thinking level, combining reinforcement learning algorithms and user characteristics, replies that match deep with the thinking level are generated, which solves the problem of insufficient flexibility in answering questions by artificial intelligence models, and achieves more flexible and easy-to-understand reply information generation.

CN120354949APending Publication Date: 2025-07-22LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510551021.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing artificial intelligence models are less flexible when answering questions, with fixed and single patterns, making it difficult to provide flexible response information according to different user needs.

Method used

By obtaining target parameters that represent the depth of thinking level, the target model is used to generate response information that matches the depth of thinking level, including multiple thinking levels of different depths, the value table is updated using reinforcement learning algorithms, and combined with user identity characteristics and problem areas, the dialogue is simulated to determine the reply information.

Benefits of technology

It improves the flexibility and adaptability of the artificial intelligence model in answering questions, and can provide richer and easier to understand response information according to user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354949A_ABST
    Figure CN120354949A_ABST
Patent Text Reader

Abstract

The invention discloses an information processing method and electronic equipment. The method comprises the steps of obtaining a target question; a target parameter representing the depth of the thinking level is obtained, the thinking level comprises multiple thinking levels of different depths, and the target parameter is used for representing one of the multiple thinking levels of different depths; based on the target parameters and the target model, reply information for the target question is obtained, and the reply information is deeply matched with the thinking level represented by the target parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to an information processing method and an electronic device. Background Art

[0002] With the continuous development of artificial intelligence technology, application scenarios of using artificial intelligence models such as large language models to answer questions raised by users are becoming increasingly widespread.

[0003] However, the modes of answering questions by means of artificial intelligence models are fixed and single, with poor flexibility. Summary of the Invention

[0004] On the one hand, this application provides an information processing method, including:

[0005] Obtain a target question;

[0006] Obtain a target parameter characterizing the depth of the thinking level, where the thinking level includes multiple thinking levels with different depths, and the target parameter is used to characterize one of the multiple thinking levels with different depths;

[0007] Based on the target parameter and a target model, obtain a reply message for the target question, where the reply message matches the depth of the thinking level characterized by the target parameter.

[0008] In a possible implementation manner, the obtaining of the target parameter characterizing the depth of the thinking level includes:

[0009] Display different depth options for the thinking level;

[0010] In response to a selection operation, obtain a target parameter characterizing the depth of the thinking level, where the selection operation is a selection of one of the different depth options for the thinking level.

[0011] In another possible implementation manner, the obtaining of the target parameter characterizing the depth of the thinking level includes:

[0012] Based on the target question, obtain a target parameter characterizing the depth of the thinking level;

[0013] Wherein, the target question belongs to one of multiple recommended questions, and the multiple recommended questions are predicted questions generated and displayed in the previous conversation.

[0014] In another possible implementation manner, different depths of the thinking level correspond to different numbers of dialogue rounds of the simulated conversation;

[0015] The obtaining of the reply message for the target question based on the target parameter and the target model includes:

[0016] Based on the target number of dialogue turns indicated by the target parameter, perform a simulated dialogue on the target question through a target model to obtain response information for the target question.

[0017] In another possible implementation, the performing a simulated dialogue on the target question through a target model based on the target number of dialogue turns indicated by the target parameter includes:

[0018] Based on the target number of dialogue turns indicated by the target parameter, perform a simulated dialogue on the target question through at least two agents, where the first agent is used to generate an answer through a first model, and the second agent is used to generate a predicted question through a second model, and both the first agent and the second agent belong to the at least two agents;

[0019] Obtain the target answer obtained by the first agent and the target predicted question obtained by the second agent in the last round of the simulated dialogue;

[0020] Determine the response information based on the target answer and the target predicted question.

[0021] In another possible implementation, the performing a simulated dialogue on the target question through a target model based on the target number of dialogue turns indicated by the target parameter includes:

[0022] Based on the target number of dialogue turns indicated by the target parameter, perform the simulated dialogue of the target number of dialogue turns on the target question through the first agent and the second agent, where the first agent is used to generate an answer through a first model, and the second agent is used to generate a predicted question through a second model;

[0023] For each round of simulated dialogue through the first agent and the second agent, obtain the current round answer obtained by the first agent in this round of simulated dialogue and the current round predicted question obtained by the second agent in this round of simulated dialogue, and determine the relevance between the current round answer and the current round predicted question and the target question through a third agent. If the relevance is less than a set threshold, terminate the simulated dialogue between the first agent and the second agent, and the third agent determines the relevance through a third model;

[0024] In response to the end of the simulated dialogue between the first agent and the second agent, use a fourth agent to determine the response information from at least one round of answers and at least one round of questions obtained from the simulated dialogue, and the fourth agent determines the response information through a fourth model.

[0025] In another possible implementation, the obtaining a target parameter characterizing the depth of the thinking level based on the target question includes:

[0026] Determine the historical thinking level depth corresponding to the target question based on the historical thinking level depths adopted by each of the multiple recommended questions generated based on the previous session.

[0027] Update the value table using a reinforcement learning algorithm based on the reward score associated with the historical thinking level depth corresponding to the target question and the problem domain corresponding to the target question, where the value table is a relationship table between different problem domains and parameters representing different thinking level depths.

[0028] Determine the target parameter representing the thinking level depth based on the problem domain corresponding to the target question and in combination with the updated value table.

[0029] In another possible implementation, the updating the value table using a reinforcement learning algorithm based on the reward score associated with the historical thinking level depth corresponding to the target question and the problem domain corresponding to the target question includes:

[0030] Obtain the user identity characteristics of the user.

[0031] Update the value table using a reinforcement learning algorithm based on the reward score associated with the historical thinking level depth corresponding to the target question, the problem domain corresponding to the target question, and the user identity characteristics, where the value table is a relationship table between different problem domains and user attribute characteristics and parameters representing different thinking level depths.

[0032] The determining the target parameter representing the thinking level depth based on the problem domain corresponding to the target question and in combination with the updated value table includes:

[0033] Determine the target parameter representing the thinking level depth based on the problem domain corresponding to the target question and the user identity characteristics and in combination with the updated value table.

[0034] In another possible implementation, the obtaining the reply information for the target question based on the target parameter and the target model includes:

[0035] Determine a target model that matches the thinking level depth represented by the target parameter from among multiple models corresponding to different thinking level depths.

[0036] Obtain the reply information for the target question using the target model.

[0037] In another aspect, the present application also provides an information processing device, including:

[0038] A problem obtaining unit, configured to obtain a target question.

[0039] A depth determination unit for obtaining a target parameter characterizing the depth of a thinking level, where the thinking level includes multiple thinking levels with different depths, and the target parameter is used to characterize one of the multiple thinking levels with different depths;

[0040] A reply obtaining unit for obtaining reply information for the target question based on the target parameter and a target model, where the reply information matches the depth of the thinking level characterized by the target parameter.

[0041] In another aspect, the present application also provides an electronic device, including:

[0042] An output device;

[0043] A processor for being in a working state based on an intelligent program. The intelligent program obtains a target question and a target parameter characterizing the depth of a thinking level, where the thinking level includes multiple thinking levels with different depths, and the target parameter is used to characterize one of the multiple thinking levels with different depths; the intelligent program provides the target parameter and the target question to a target model, and the intelligent program has the ability to call the target model; the intelligent program obtains the reply information of the target model for the target question and outputs the reply information through the output device, and the reply information matches the depth of the thinking level characterized by the target parameter. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.

[0045] Figure 1 It is a schematic flowchart of an information processing method provided by the present application;

[0046] Figure 2 It is a schematic diagram of a session interaction interface provided by the present application;

[0047] Figure 3 It is another schematic flowchart of an information processing method provided by the present application;

[0048] Figure 4 It is another schematic flowchart of an information processing method provided by the present application;

[0049] Figure 5 It is another schematic flowchart of an information processing method provided by the present application;

[0050] Figure 6 It is an example diagram of an implementation framework of the information processing method in the present application;

[0051] Figure 7 A schematic diagram of an implementation process of another information processing method provided for this application;

[0052] Figure 8 A schematic diagram of a composition structure of an information processing device provided for this application;

[0053] Figure 9 A schematic diagram of a composition architecture of an electronic device provided for this application. Detailed implementation manners

[0054] The embodiments of this application will be described below in conjunction with the accompanying drawings in the embodiments of this application. The terms used in the implementation manners part of this application are only used to explain the specific embodiments of this application, rather than aiming to limit this application. Those of ordinary skill in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0055] The terms "first", "second", etc. in the specification, claims and above-mentioned accompanying drawings of this application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinction adopted when describing objects with the same attributes in the embodiments of this application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0056] The target model involved in the embodiments of this application is a machine learning model, which can identify natural language and / or other inputs (such as audio and video, images, tables, etc.) input to the target model, and perform comprehensive language processing tasks such as semantic analysis and answering questions, and then generate outputs related to the input and / or respond to the input.

[0057] The target model involved in the embodiments of this application learns the features and rules of natural language by training a large number of diverse data, so as to be able to understand and generate natural language. It usually has model parameters in the hundreds of millions to hundreds of billions (model parameters are variables that control the behavior of the target model), and can capture complex relationships and patterns in natural language.

[0058] In the embodiments of the present application, the target model involved may be a generative model or a generative language model (GLMs). For example, it may specifically include large language models (LLMs), GPT (Generative Pre-trained Transformer), etc. The model involved in the embodiments of the present application may be a general large model or an expert large model obtained by fine-tuning based on requirements. The embodiments of the present application do not make any limitations in this regard.

[0059] As Figure 1 , a schematic flowchart of an information processing method provided by the present application is shown. The method of this embodiment can be applied to an electronic device, which can be a device node in a system such as a server or a cloud platform, or a terminal device such as a laptop or a desktop computer, without limitation.

[0060] The method of this embodiment may include:

[0061] S101, obtain a target problem.

[0062] Wherein, the target problem is a problem that needs to be answered.

[0063] For example, the target problem may be a problem input by the user.

[0064] For another example, considering that after each round of conversation, the electronic device can also predict at least one recommended question based on an intelligent model, etc. The recommended question is a predicted question that the user may wish to follow up. On this basis, the target problem may also be a question selected by the user from at least one recommended question generated and displayed in the previous round of conversation.

[0065] S102, obtain a target parameter characterizing the depth of the thinking level.

[0066] In this application, the thinking levels may include multiple thinking levels with different depths. The depth of a thinking level represents the degree of thinking or the depth of knowledge search for answering a question based on an intelligent model (such as a target model, etc.), and also represents the comprehensibility or the richness of the content of the answer given by the intelligent model. Based on this, the depth of the thinking level is related to the amount of knowledge to be mined for the intelligent model to perform reasoning, the number of dimensions of information, the types of disciplines to be analyzed, the total number of types of disciplines for interdisciplinary analysis, and the angles for analyzing problems during the reasoning process, etc. Among them, the deeper the depth of the thinking level, the more knowledge the intelligent model needs to mine for reasoning, the more information dimensions are covered, the greater the knowledge difficulty corresponding to the disciplines being analyzed, the more types of disciplines for interdisciplinary analysis, and the more angles for analyzing problems. Based on this, the deeper the depth of the thinking level, the more comprehensive the information in the answer given to the question, and the higher the difficulty of understanding, etc.

[0067] Among them, the target parameter is used to represent one of the multiple thinking levels with different depths. For example, the target parameter can be a depth level, a depth score, a character, or a numerical value representing the depth of the thinking level, etc., without specific limitations.

[0068] It can be understood that the order of step S102 and step S101 is not limited to Figure 1 As shown, in practical applications, it may also be that the target parameter is obtained while obtaining the target question.

[0069] S103, based on the target parameter and the target model, obtain a response message for the target question.

[0070] Among them, the target model is an intelligent model that gives a response to the target question. The target model can include a single intelligent model or may be an integration of multiple intelligent models, without any limitations in this regard.

[0071] In this application, the response message matches the depth of the thinking level represented by the target parameter. That is to say, the response message is the response obtained after reasoning at the depth of the thinking level represented by the target parameter for the target question. It can be seen from this that for the same target question, when the depth of the thinking level represented by the target parameter is different, the amount of information and the degree of understandability, etc. of the response message obtained based on the target model will also vary.

[0072] In this application, the response message at least includes the answer given to the target question, and may also include at least one recommended question for the user to choose from for the target question. Among them, the recommended question in the response message is a question predicted for the target question that the user may ask.

[0073] As can be seen from the above, this application does not directly perform reasoning processing on the target problem. Instead, it obtains a target parameter representing the depth of the thinking level, and based on the target model, obtains a response message for the target problem that matches the depth of the thinking level represented by the target parameter. Therefore, when the depth of the thinking level represented by the target parameter is different, the generated response message will also be different, so that corresponding response messages can be generated according to different requirements for the depth of the thinking level, improving the flexibility of the response message given for the target problem.

[0074] In this application, there can be multiple possible ways to obtain the target parameter. Several possible situations are described below.

[0075] In a possible situation of obtaining the target parameter, this application can display different depth options for the thinking level. Correspondingly, in response to a selection operation, a target parameter representing the depth of the thinking level is obtained. The selection operation is a selection of one of the different depth options for the thinking level.

[0076] Among them, different depth options represent different depths of the thinking level. Correspondingly, after selecting a target depth option from the different depth options through the selection operation, the target parameter corresponding to the depth of the thinking level corresponding to the target depth option can be determined.

[0077] For example, assume that the numbers 1 - 10 are used to represent the thinking level depths from the first level to the tenth level respectively, where the larger the level number, the deeper the thinking level. If the target depth option selected by the user corresponds to the fifth-level thinking level depth, then the target parameter can be the number 5 corresponding to the fifth-level thinking level depth.

[0078] It can be understood that for the situation where the user inputs a target problem, or the user selects a target problem from the recommended problems predicted in the previous round, etc., this possible situation can be used to determine the target parameter, without specific restrictions.

[0079] For ease of understanding, reference can be made to Figure 2 , Figure 2 which shows a schematic diagram of the session interaction interface provided by this application.

[0080] From Figure 2 it can be seen that the session interaction interface includes a session message interaction area 201 and a question input box 202.

[0081] Among them, in the session interaction area 201, in addition to displaying the questions raised by the user and the answers determined for the questions based on the target model, there is also a recommended question display bar 203, which can display at least one recommended question predicted in the previous round of the session. Of course, Figure 2This is just an example. In actual applications, the recommended question display bar may not be shown, or it may be shown in other areas outside the session interaction area. There are no specific restrictions.

[0082] The question input box 202 can be used to input questions that the user needs to answer. In this application, multiple depth options 204 that can be selected by the user are also displayed in the question input box, such as Figure 2 In, the options corresponding to depth level 1 to depth level 4 are all depth options.

[0083] In Figure 2 In, different depth options represent different levels of thinking depth. For example, in Figure 2 Taking the depth options as depth level 1, depth level 2, depth level 2, and depth level 4 respectively as an example. Among them, the thinking level corresponding to depth level 4 has the deepest depth, while the thinking level corresponding to thinking level 1 has the shallowest depth.

[0084] Of course, there may be other possibilities for the depth options. For example, in order to more intuitively reflect the users to whom the depth of the thinking level applies, the depth options can also include: applicable to ordinary adults, representing the ordinary depth level with a relatively deep thinking level; applicable to highly educated adults, representing the highly knowledgeable depth level with the deepest thinking level; applicable to the elderly, representing the medium depth level with a moderate depth of the thinking level; and applicable to children, representing the children's depth level with a relatively shallow depth of the thinking level.

[0085] In Figure 2 In the shown session interaction interface, if the user hopes to manually input a question, then the user can input the target question to be answered in the question input box. The user can trigger the option or icon of the target question by clicking to confirm, etc., to achieve the on-screen output of the target question, so that the electronic device can obtain the target question. Before or after the on-screen output of the target question, the user can also select a depth option from multiple depth options, so that the electronic device can determine the target parameter corresponding to the depth option selected by the user. On this basis, as Figure 2 Shown, the electronic device can determine the corresponding reply information suitable for the target question input by the user based on the target parameter with the help of the target model.

[0086] If the user believes that there is a question that the user hopes to ask among the recommended questions predicted in the previous round of the session, then the user can also directly select the target question from the recommended questions, and then the electronic device can obtain the target question selected by the user. Among them, before or after the user selects the target question from the recommended questions, the user can also select the required depth option from multiple depth options, so that the electronic device can obtain the target parameter corresponding to the depth option selected by the user.

[0087] It should be noted that Figure 2 This is just an example of the conversation interaction interface. In actual applications, the recommended question display bar can also be in other areas outside the conversation interaction area. Correspondingly, the depth option can also be displayed in other areas outside the question input box, and there is no restriction on this.

[0088] In another possible case of obtaining the target parameter, the present application can obtain a target parameter characterizing the depth of the thinking level based on the target question.

[0089] Among them, there can be multiple possible implementation manners for obtaining the target parameter based on the target question.

[0090] For example, in the first possible implementation manner, the problem domain corresponding to the target question can be determined, and based on the problem domain corresponding to the target question, a target parameter characterizing the depth of the thinking level can be obtained.

[0091] Among them, the problem domain of the target question can characterize the knowledge category or problem background related to the target question, etc. There can be multiple classification rules for the problem domain. For example, the problem domain can be divided into multiple main problem domains such as the scientific domain, the technical domain, the art domain, and the social domain. Each main problem domain can be further divided into different sub-problem domains. For example, the scientific domain can be divided into the physics domain, the chemistry domain, the mathematics domain, and the biology domain, etc.; while the social domain can be divided into the economic-related domain, the culture-related domain, and the daily life domain, etc.

[0092] In the present application, there can be multiple possible implementations for determining the problem domain corresponding to the target question. For example, based on the target question and at least one domain keyword corresponding to each of the multiple pre-divided problem domains, the problem domain with the domain keyword matching the target question can be determined, and the matched problem domain can be determined as the problem domain corresponding to the target question. Another example is that the problem domain corresponding to the target question can be determined by using a domain classification model, and the domain classification model can be pre-trained based on different problem samples and the actual problem domains annotated for the problem samples. The domain classification model can be a neural network model or other intelligent models, etc., and there is no restriction on this.

[0093] On this basis, there can also be multiple implementations for determining the target parameter matching the problem domain corresponding to the target question. For example, according to actual needs, depth parameters (also simply referred to as parameters) suitable for different problem domains can be pre-configured. The depth parameter is used to characterize the depth of the thinking level. Correspondingly, the depth parameter matching the problem domain corresponding to the target question can be queried, and the matched depth parameter can be determined as the target parameter.

[0094] For another example, the present application can also combine a reinforcement learning algorithm to construct a value table between different problem domains and parameters representing different levels of thinking depth. The value table is a relationship table between different problem domains and different parameters. Of course, the parameters in the value table can also be the levels of thinking depth. On this basis, the present application can combine the value table to determine the target parameters matching the problem domain corresponding to the target problem. Among them, the value table can be updated regularly or irregularly.

[0095] In a second possible implementation manner, the target parameters representing the thinking level depth can be obtained based on the problem domain corresponding to the target problem and the user identity characteristics of the user corresponding to the target problem. For example, the mapping relationship between different problem domains, different user identity characteristics, and different parameters representing the thinking level depth can be pre-configured, and based on this mapping relationship, the target parameters matching the problem domain corresponding to the target problem and the user identity characteristics of the user can be determined.

[0096] It can be understood that the above two possible implementation manners are not only suitable for the scenario where the target problem is a problem manually input by the user, but also suitable for the scenario where the target problem comes from multiple recommended problems.

[0097] In a third possible implementation manner, on the premise that the target problem belongs to one of the multiple predicted problems generated and displayed in the previous session, the present application can also combine the reinforcement learning algorithm to first update the value table, and then determine the target parameters based on the updated value table and the problem domain of the target problem. The following combines Figure 3 to illustrate this implementation manner.

[0098] Such as Figure 3 shows another schematic flowchart of the information processing method provided by the present application. The method of this embodiment may include:

[0099] S301, obtain the target problem.

[0100] Among them, the target problem belongs to one of the multiple recommended problems. The multiple recommended problems are the predicted problems generated and displayed in the previous session.

[0101] Among them, the previous session refers to the session for the most recent problem before the target problem. It can be understood that before obtaining the target problem, for the most recent problem, the electronic device can generate an answer to the most recent problem and multiple recommended problems.

[0102] S302, based on the historical thinking level depth adopted by each of the multiple recommended problems generated in the previous session, determine the historical thinking level depth corresponding to the target problem.

[0103] It can be understood that in the previous conversation, different recommended questions are generated by adopting different depths of thinking levels. Therefore, different recommended questions correspond to different depths of thinking levels. In order to distinguish from the depth of thinking level involved in the current reasoning target question, the depth of thinking level corresponding to the recommended question generated in the previous conversation is called the historical depth of thinking level.

[0104] It can be understood that since the target question belongs to multiple recommended questions, therefore, when the historical depths of thinking levels corresponding to multiple recommended questions are determined, the historical depth of thinking level corresponding to this target question can be obtained.

[0105] S303, based on the reward score associated with the historical depth of thinking level corresponding to the target question and the problem domain corresponding to the target question, use the reinforcement learning algorithm to update the value table.

[0106] Among them, the value table is a relationship table between different problem domains and parameters representing different depths of thinking levels. For example, the value table can include multiple problem domains and the values of different parameters corresponding to each problem domain. The value of each parameter corresponding to a problem domain represents the suitability of this problem domain for the depth of thinking level represented by this parameter. Each parameter represents a different depth of thinking level.

[0107] Among them, the reward score associated with the historical depth of thinking level corresponding to this target question can be determined according to the set reward mechanism. This reward mechanism can be set according to actual needs.

[0108] For example, the reward mechanism is: determine the reference depth of thinking level actually used to infer the answer in the previous conversation, and based on the depth gap between the historical depth of thinking level corresponding to the target question and this reference depth of thinking level, determine the reward score corresponding to this depth gap. Among them, the depth gap can be the number of depth levels or the depth value difference between the historical depth of thinking level corresponding to the target question and this reference depth of thinking level, etc. Among them, different depth gaps or different ranges of depth gaps correspond to different reward scores.

[0109] Another example, the reward mechanism can also be: if the historical depth of thinking level corresponding to the target question is this reference depth of thinking level, the reward score is the first reward score; if the historical depth of thinking level corresponding to this target question is not this reference depth of thinking level, the reward score is the second reward score, and this first reward score is greater than this second reward score.

[0110] Among them, the specific implementation process of using the reinforcement learning algorithm to update the value table can be unrestricted.

[0111] For example, for the sake of easy understanding, taking the depth value of the parameter in the value table as the depth of thinking level as an example, the value table is a matrix, and the size of this matrix is , where is the number of types of problem domains, represents the set maximum depth of the thinking level, is the set minimum depth of the thinking level. Based on this, each problem domain is a state in reinforcement learning, and the depth of the thinking level is an action in reinforcement learning. Then the state sequences corresponding to types of problem domains, where represents the th type of problem domain, is an integer from 1 to . Then reinforcement learning can be performed in combination with the Bellman equation shown in the following formula to update this value table :

[0112]

[0113] where represents the depth of the thinking level corresponding to , represents the set learning rate, is the set decay coefficient, represents the maximum value that can be obtained in state , , represents the depth of the thinking level corresponding to .

[0114] S304, based on the problem domain corresponding to the target problem, in combination with the updated value table, determine the target parameter representing the depth of the thinking level.

[0115] For example, from the updated value table, select the parameter with the highest value and related to the problem domain corresponding to the target problem as the target parameter.

[0116] Of course, in combination with the value scores of the problem domain corresponding to the target problem in the updated value table on different parameters, specific strategies can also be considered to determine the target parameter, without specific limitations.

[0117] S305, based on the target parameter and the target model, obtain the response information for the target problem.

[0118] where the response information matches the depth of the thinking level represented by the target parameter.

[0119] This step S305 can refer to the relevant introduction in other embodiments of this application and will not be elaborated here.

[0120] It can be understood that inFigure 3 Take the example where the value table is only related to the problem domain. However, in practical applications, considering the different preferences and educational levels of different users, even for the same problem, the depth levels of the responses expected by different users will vary.

[0121] Based on this, in a possible implementation manner, the user identity characteristics of the user can also be obtained in this application. Correspondingly, the value table can be updated using a reinforcement learning algorithm based on the reward score deeply associated with the historical thinking level corresponding to the target problem, the problem domain corresponding to the target problem, and the user identity characteristics of this user.

[0122] In this implementation manner, the value table is a relationship table between different problem domains and user attribute characteristics and parameters representing different thinking level depths. Among them, the value tables corresponding to different users will also be different, enabling the value table to reflect what different users expect for problems in different domains.

[0123] Among them, the user for whom the user identity characteristics need to be obtained here can be the user of this electronic device, that is, the user who selects or enters this target problem.

[0124] For example, based on the user account of the logged-in user in the session, the user identity characteristics of the user can be determined. For example, obtain the user identity characteristics associated with this user account. The user identity characteristics can be identity characteristic information such as the user's age, educational level, and hobbies that can be obtained with the user's permission.

[0125] For another example, the user identity characteristics of the user can be determined by combining the problems proposed in the user's historical session, etc., without specific limitations.

[0126] In this implementation manner, different user identity characteristics and different problem domains in the value table represent different states, while different parameters represent different actions. Therefore, except that the states in the value table are different from the states in the previous value table, the specific implementation process of updating the value table is similar and will not be elaborated here.

[0127] Correspondingly, after updating the value table, based on the problem domain corresponding to the target problem and the user identity characteristics of this user, combined with the updated value table, determine the target parameter representing the thinking level depth. For example, based on the problem domain corresponding to the target problem and the user identity characteristics, query the target parameter in the value table that best matches this problem domain and the given identity characteristic information.

[0128] In any of the above embodiments of this application, there can be multiple possible specific implementations for obtaining the response information that matches the thinking level depth represented by the target parameter. The following several possible situations will be described.

[0129] In a possible scenario, the present application can pre-train multiple models with different depths of thinking levels. After each model is trained, the depth of the thinking level corresponding to each model is fixed. Therefore, the depth of the thinking level adopted by each model to process problems is also fixed. Among them, models with different depths of thinking levels can be models of the same type but with different internal parameters, or models of different types, without any restrictions in this regard. Among them, different models have different thinking levels when performing task reasoning. Correspondingly, for the same problem, the information depth and understanding difficulty of the reply information such as the answers inferred by different models will also vary.

[0130] Based on this, the present application can determine a target model that matches the depth of the thinking level represented by the target parameter from multiple models corresponding to different depths of thinking levels. Correspondingly, the reply information for the target problem can be obtained using the target model, so that the target model can undergo reasoning processing that matches the depth of the thinking level corresponding to the target parameter to obtain reply information that matches the depth of the thinking level represented by the target parameter.

[0131] In another possible scenario, different depths of thinking levels can correspond to different numbers of dialogue turns in a simulated dialogue. Among them, a simulated dialogue refers to a way of inferring the reply information of a problem at least by simulating the dialogue between a questioner and a responder. It can be understood that since each round of simulated dialogue requires the model to perform reasoning, as the number of rounds of simulated dialogue increases, the amount of reasoning required by the model increases, that is, the depth of the thinking level required to obtain the problem increases.

[0132] Based on this, the present application can perform a simulated dialogue on the target problem through the target model based on the target number of dialogue turns indicated by the target parameter to obtain the reply information for the target problem.

[0133] Among them, the number of dialogue times for the target model to perform a simulated dialogue on the target problem is the target number of dialogue turns.

[0134] Among them, there can be various specific implementations for the target model to perform a simulated dialogue on the target problem, and the present application does not impose any restrictions in this regard.

[0135] The following will illustrate several possible implementation methods for the target model to perform a simulated dialogue on the target problem.

[0136] In a possible implementation method, it can be as Figure 4 , Figure 4 shows another flowchart of the information processing method provided by the present application. The method of this embodiment can include:

[0137] S401, obtain the target problem.

[0138] S402. Obtain the target parameter characterizing the depth of the thinking level.

[0139] The thinking levels in this application include multiple thinking levels with different depths.

[0140] Among them, the target parameter is used to characterize one of the multiple thinking levels with different depths.

[0141] In this embodiment, the thinking levels with different depths correspond to different numbers of dialogue turns in the simulated dialogue.

[0142] For the above steps S401 and S402, reference can be made to the relevant introductions in other embodiments of this application, which will not be elaborated here.

[0143] S403. Based on the target number of dialogue turns indicated by the target parameter, perform a simulated dialogue on the target question through at least two agents.

[0144] Among them, the target number of dialogue turns is the number of dialogue turns corresponding to the depth of the thinking level characterized by the target parameter. The target number of dialogue turns represents the number of turns required to perform a simulated dialogue on the target question.

[0145] Among them, an agent can be a program or system for the user role in the simulated dialogue. In this application, the agent can simulate the user role to participate in the dialogue through a model (any artificial intelligence model).

[0146] In this embodiment, the at least two agents include: a first agent and a second agent. Among them, the first agent is used to generate an answer through a first model. The second agent is used to generate a predicted question through a second model.

[0147] For example, using the target question as the input question for the first round of simulated dialogue, perform the simulated dialogue the target number of dialogue turns through the first agent and the second agent. Among them, in each round of simulated dialogue, the first agent generates the answer for this round corresponding to the input question of this round of simulated dialogue based on the first model; the second agent determines the predicted question for this round corresponding to the input question of this round of simulated dialogue based on the second model, and uses the predicted question for this round as the input question for the next round of simulated dialogue to perform the next round of simulated dialogue.

[0148] Among them, the first agent generating the answer for this round corresponding to the input question of this round of simulated dialogue based on the first model can be that the first agent generates the answer for this round only based on the input question of this round of simulated dialogue and based on the first model.

[0149] Specifically, if the current simulated conversation is not the first simulated conversation, the first intelligent agent can also generate the current answer based on the input question of the current simulated conversation and the previous answer obtained from the previous simulated conversation, using the first model. For example, the first model adjusts the previous answer in combination with the input question, such as correcting the previous answer and filling in information related to the current input question to enrich and improve the content of the answer, etc.

[0150] When the second intelligent agent generates the current predicted question based on the second model, it can be that the second intelligent agent generates the predicted question only based on the input question of the current simulated conversation. For example, the second intelligent agent constructs a first prompt based on the input question of the current round and inputs the first prompt into the second model to obtain the predicted question generated by the second model. Here, the first prompt is used to prompt the large language model to predict the question that the user may ask next based on this input question.

[0151] Another way for the second intelligent agent to generate the current predicted question based on the second model is that the second intelligent agent generates the predicted model based on the input question of the current simulated conversation and the current answer, using the second model. For example, the second intelligent agent generates a second prompt based on the input question of the current round and the current answer. The second prompt is used to prompt the large language model to predict the question that the user may ask after this input question based on this input question and the current answer. For example, the first prompt can be "The goal is to predict the question that the user may ask next based on the conversation content between you and the user. The conversation content is: the input question of the current round, the current answer".

[0152] It can be understood that the at least two intelligent agents may also include an optimization intelligent agent, which is used to optimize and adjust the answer generated by the first intelligent agent through the first model based on the optimization model. For example, the optimization intelligent agent optimizes the answer generated by the first model based on the user identity characteristics using the optimization model, so that the optimized answer better conforms to the user identity characteristics. For example, the semantics and description method of the optimized answer are more in line with the user identity characteristics.

[0153] It should be noted that in this embodiment, the first model used by the first intelligent agent and the second model used by the second intelligent agent can be different models. The first model and the second model can also be the same model, but the first intelligent agent and the second intelligent agent instruct the model to perform different task processes.

[0154] S404, obtain the target answer obtained by the first intelligent agent and the target predicted question obtained by the second intelligent agent in the last round of simulated conversation.

[0155] It can be understood that for each additional round of the simulated conversation, the level of thinking used to obtain the answer and predict the question deepens by one level. Therefore, the answer obtained in the last round of the simulated conversation is the answer corresponding to the depth of the thinking level characterized by the target parameter. Correspondingly, the predicted question obtained in the last round of the simulated conversation is the predicted question corresponding to the depth of the thinking level characterized by the target parameter.

[0156] For the sake of easy distinction, the answer obtained in the last round of the simulated conversation is referred to as the target answer, and the predicted question obtained in the last round of the simulated conversation is referred to as the target predicted question.

[0157] S405. Determine the reply information based on the target answer and the target predicted question.

[0158] Among them, the reply information includes at least the target answer, and may also include the target answer and the target predicted question at the same time.

[0159] For example, the target answer can be used as the reply information and output to the conversation interaction area of the conversation interaction window, so that the user can see the target answer to the target question in the conversation interaction area. On this basis, the target predicted question can be separately output as the predicted question for this round of conversation corresponding to the target question to the predicted question display area. Of course, while outputting the target predicted question as a recommended question, the predicted questions obtained in each round of the simulated conversation in the target conversation round can also be used as recommended questions and output each recommended question.

[0160] For another example, since the target predicted question is the predicted question generated in the most recent round of the simulated conversation, this target predicted question is also the question that the user is most likely to want to ask. On this basis, the present application can also use both the target answer and the target predicted question as the reply information. Correspondingly, after outputting the reply information, the user can not only see the target answer, but also more intuitively see the recommendation of the target predicted question. Among them, outputting the reply information can be outputting the reply information in the conversation interaction area. While outputting the reply information, the present application can also display the predicted questions obtained in each round of the simulated conversation as recommended questions in other areas outside the conversation interaction area in the conversation interaction window, so that the user can more conveniently select the recommended questions that may be desired to input.

[0161] The following combines Figure 5 Another implementation manner of the target model for simulating a conversation for the target question is introduced. For example Figure 5 shows another schematic flowchart of the information processing method provided by the present application. The method of this embodiment may include:

[0162] S501. Obtain the target question.

[0163] S502. Obtain a target parameter characterizing the depth of the thinking level.

[0164] The thinking levels in this application include multiple thinking levels with different depths.

[0165] Among them, the target parameter is used to characterize one of the multiple thinking levels with different depths.

[0166] In this embodiment, the thinking levels with different depths correspond to different numbers of dialogue turns in the simulated dialogue.

[0167] For the above steps S501 and S502, reference can be made to the relevant introductions in the previous embodiments, which will not be elaborated here.

[0168] S503. Based on the target number of dialogue turns indicated by the target parameter, conduct a simulated dialogue with the target number of dialogue turns between the first intelligent agent and the second intelligent agent for the target problem.

[0169] Among them, the first intelligent agent is used to generate an answer through the first model, and the second intelligent agent is used to generate a predicted question through the second model.

[0170] In this embodiment, the specific implementation of each round of simulated dialogue through the first intelligent agent and the second intelligent agent can be referred to the relevant introduction in the previous step S503, which will not be elaborated here.

[0171] S504. Every time a round of simulated dialogue is conducted between the first intelligent agent and the second intelligent agent, obtain the current round answer obtained by the first intelligent agent in this round of simulated dialogue and the current round predicted question obtained by the second intelligent agent in this round of simulated dialogue. Determine the relevance between the current round answer and the current round predicted question and the target problem through the third intelligent agent. If the relevance is less than the set threshold, terminate the simulated dialogue between the first intelligent agent and the second intelligent agent.

[0172] Among them, the third intelligent agent determines the relevance through the third model. For example, the third intelligent agent can generate a third prompt based on the current round answer and the current round predicted question, input the third prompt into the third model, and obtain the relevance output by the third model. Among them, the third prompt is used to prompt the relevance between the current round answer and the current round predicted question and the initial target problem. For example, the third prompt can be "Evaluate the relevance between the current round answer and the current round predicted question and the target problem. When the relevance is not large, output 0; when the relevance is large, output 1. Simulated dialogue content: current round answer, current predicted question, target problem".

[0173] It can be understood that if the relevance determined by the third intelligent agent is less than the set threshold, the simulated conversation between the first intelligent agent and the second intelligent agent will be terminated, ending the simulated conversation between the first intelligent agent and the second intelligent agent, and thus entering step S505. If the relevance determined by the third intelligent agent is not less than the set threshold, the first intelligent agent and the second intelligent agent can continue to execute the simulated conversation until the simulated conversation reaches the target number of dialogue rounds, ending the simulated conversation between the first intelligent agent and the second intelligent agent.

[0174] S505. In response to the end of the simulated conversation between the first intelligent agent and the second intelligent agent, use the fourth intelligent agent to determine the reply information from at least one round of answers and at least one round of questions obtained from the simulated conversation.

[0175] Among them, the fourth intelligent agent determines the reply information through the fourth model. Similar to the previous embodiments, the reply information at least includes: the output answer to the target question determined based on the at least one round of answers and at least one round of questions, and may also include at least one recommended question determined for the target question.

[0176] Among them, there are multiple possibilities for the fourth intelligent agent to determine the reply information through the fourth model. For example, in one possible implementation, the fourth intelligent agent generates an output answer for replying to the target question based on the at least one round of answers and at least one round of questions through the fourth model. The output answer is a newly generated answer based on the at least one round of answers and at least one round of questions. Further, through the fourth model, at least one optimization process such as question content adjustment and duplicate removal can be performed on the at least one round of questions to obtain the optimized at least one round of questions, and the optimized at least one round of questions are output as the reply information, or are output as at least one recommended question separate from the reply information.

[0177] In another possible implementation, the fourth intelligent agent generates an output answer for replying to the target question based on the at least one round of answers through the fourth model. For example, answer duplicate removal and content summary are performed on the at least one round of answers to obtain the output answer, and the output answer is used as the reply information, or the output answer that best matches the target question is selected from the at least one round of answers. Further, the fourth intelligent agent can also perform at least one optimization process such as question content adjustment and duplicate removal on the at least one round of questions to obtain at least one recommended question, as specifically described in the above implementation, and will not be elaborated here.

[0178] To facilitate understanding of the specific implementation of this embodiment, the depth of the thinking level characterized by the target parameter is taken as an example of level 2, and in combination with Figure 6 this embodiment will be introduced.

[0179] In Figure 6In the example diagram of the implementation framework shown, it includes the four agents mentioned in the above embodiments. For ease of description, these five agents are sequentially referred to as Agent A, Agent B, Agent C, and Agent D. Among them, Agent A represents the first agent for generating answers, Agent B represents the second agent for generating predicted answers, Agent C represents the third agent for determining relevance, and Agent D represents the agent for determining the problem domain. This Agent D determines the problem domain of the question based on the domain classification model.

[0180] After obtaining the question Q0, Agent D can determine the problem domain S1 corresponding to the question Q0. This problem domain S1 is essentially a type of state category when optimizing the value table using reinforcement learning.

[0181] On this basis, if the user selects a depth option from the at least one depth option displayed, the depth of the thinking level can be directly determined as the depth option selected by the user. For example, Figure 2 in the depth of the thinking level selected by the user is taken as 2 as an example.

[0182] If the user does not select a depth option and the question Q0 is the question input by the user in the first round of conversation, then the corresponding target parameters can be queried from the value table based on the problem domain S1 corresponding to the question Q0. If the question Q0 is a question selected from the recommended questions generated in the previous conversation, then this application can update the value table in a reinforcement learning manner based on the reward score and domain category corresponding to the question Q0, and determine the target parameters based on the updated value table. In Figure 6 in the depth of the thinking level corresponding to the queried target parameters is taken as 2 as an example for illustration.

[0183] Since the depth of the thinking level is 2, therefore, the number of rounds of simulated conversation required between Agent A and Agent B is two rounds.

[0184] Based on this, Agent A can be used to generate the answer A1 to the current question Q0. Then, based on the current question Q0 and this answer A1, Agent B can generate the predicted question Q1, thus completing the first round of simulated conversation.

[0185] After the first round of simulated conversation, it is necessary to use Agent C to determine the relevance of the answer A1 and the predicted question Q1 to the current question Q0. If this relevance is less than the set threshold, there is no need to execute the second round of simulated conversation between Agent A and Agent B. On this basis, the reply information can be determined based on the answer A1 and the predicted question Q1. For example,

[0186] Output both answer A1 and predicted question Q1 as response information. If the relevance is not less than the set threshold, then agent A and agent B need to continue to execute the second round of simulated conversation.

[0187] In the second round of simulated conversation, agent A determines answer A2 corresponding to predicted question Q1, while agent B generates predicted question Q2 corresponding to thinking level depth 2 based on predicted question Q1 and answer A2. At this point, two rounds of simulated conversation have been completed. Therefore, the response information can be determined based on answer A1, answer A2, predicted question Q1, and predicted question Q2. For example, the response information can be determined with the help of the aforementioned fourth agent. The specific way to determine the response information can be any of the aforementioned situations. For example, answer A2 can be used as the output answer corresponding to target question Q1, or an output answer can be generated by synthesizing answer A1 and answer A2.

[0188] In addition, both predicted questions Q1 and Q2 can be output as recommended questions. If the user selects predicted question Q2 as the question Q2 to be answered from predicted questions Q1 and Q2, then for question Q2, predicted questions Q1 and Q2 are the recommended questions generated in the previous conversation. For question Q2, it is also necessary to determine question domain S2 and perform similar processing to that of question Q1, which will not be elaborated here.

[0189] In a possible implementation manner, considering that different users have different ages, educational backgrounds, identities, etc., for the same question, the readability and tone of the answers expected by users will also be different. After generating the answer to the target question, the present application can also combine the user identity characteristics corresponding to the target question and convert the answer into an answer that matches the user's identity characteristics. For example Figure 7 shows a schematic flowchart of another information processing method provided by the present application. The method of this embodiment may include:

[0190] S701, obtain the target question.

[0191] This step can refer to the relevant introduction in any previous embodiment and will not be elaborated here.

[0192] S702, determine the response information based on the target question and the target model.

[0193] For example, the target question can be input into the target model to obtain the response information generated by the target model. Among them, the thinking level depth of the target model for processing the target model is fixed.

[0194] In this embodiment, the response information at least includes the answer corresponding to the target question. Of course, the response information may also include: at least one predicted question that can be recommended for the user to select.

[0195] S703. Based on the user identity characteristics corresponding to the user who proposed the target question, the target agent processes and converts the reply information to obtain reply information that matches the user identity characteristics corresponding to the user.

[0196] Among them, the specific implementation method for obtaining the user identity characteristics of the user and the specific meaning of the user identity characteristics can be referred to the relevant introduction in the previous embodiments, which will not be elaborated here.

[0197] Among them, the target agent can optimize at least one of the answer in the reply information and the prediction question based on a specific model or through a specific algorithm, etc. For example, if the user identity characteristics indicate that the user is a primary school student, then the content of the answer in the reply information can be refined or simplified so that the user can understand the specific meaning of the answer. Similar processing can also be performed on the prediction question in the reply information. If the user identity characteristics indicate that the user is a senior intellectual, then the answer in the reply information can be expanded or enriched, etc., so that the user can learn more knowledge related to the target question.

[0198] In this embodiment, after determining the reply information based on the target question, the user identity characteristics of the user who proposed the target question are also obtained, and the target agent is used to convert the reply information into reply information that matches the user identity characteristics, thereby facilitating improving the readability of the reply information for the user and reducing the situation where the user cannot understand the reply information or the amount of information obtained from the reply information is small.

[0199] On the other hand, corresponding to the information processing method provided in this application, this application also provides an information processing device. As shown in FIG. 8, a schematic diagram of a composition structure of the information processing device provided in this application is shown. The device in this embodiment may include:

[0200] A problem acquisition unit 801, configured to acquire a target question;

[0201] A depth determination unit 802, configured to acquire a target parameter characterizing the depth of the thinking level, where the thinking level includes multiple thinking levels with different depths, and the target parameter is used to characterize one of the multiple thinking levels with different depths;

[0202] A reply acquisition unit 803, configured to acquire reply information for the target question based on the target parameter and the target model, and the reply information matches the depth of the thinking level characterized by the target parameter.

[0203] In a possible implementation manner, the depth determination unit includes:

[0204] An option display subunit, configured to display different depth options for the thinking level;

[0205] A depth determination subunit, configured to obtain a target parameter characterizing the depth of the thinking level in response to a selection operation, where the selection operation is a selection of one of different depth options for the thinking level.

[0206] In another possible implementation, the depth determination unit includes:

[0207] A depth analysis subunit, configured to obtain a target parameter characterizing the depth of the thinking level based on the target question; where the target question belongs to one of multiple recommended questions, and the multiple recommended questions are prediction questions generated and displayed in the previous session.

[0208] In another possible implementation, the depth analysis subunit includes:

[0209] A level determination subunit, configured to determine the historical thinking level depth corresponding to the target question based on the historical thinking level depths adopted by the multiple recommended questions generated in the previous session;

[0210] A table update subunit, configured to update a value table using a reinforcement learning algorithm based on the reward score associated with the historical thinking level depth corresponding to the target question and the problem domain corresponding to the target question, where the value table is a relationship table between different problem domains and parameters characterizing different thinking level depths;

[0211] A parameter determination subunit, configured to determine a target parameter characterizing the depth of the thinking level based on the problem domain corresponding to the target question and in combination with the updated value table.

[0212] In another possible implementation, the table update subunit includes:

[0213] A feature acquisition subunit, configured to acquire user identity features of the user;

[0214] An update processing subunit, configured to update a value table using a reinforcement learning algorithm based on the reward score associated with the historical thinking level depth corresponding to the target question, the problem domain corresponding to the target question, and the user identity features, where the value table is a relationship table between different problem domains and user attribute features and parameters characterizing different thinking level depths;

[0215] The parameter determination subunit is specifically configured to determine a target parameter characterizing the depth of the thinking level based on the problem domain corresponding to the target question and the user identity features and in combination with the updated value table.

[0216] In another possible implementation, thinking levels of different depths correspond to different numbers of dialogue turns in a simulated dialogue;

[0217] The reply obtaining unit is specifically configured to perform a simulated conversation on the target question through a target model based on the target conversation turn number indicated by the target parameter, so as to obtain reply information for the target question.

[0218] In another possible implementation manner, the reply obtaining unit includes:

[0219] A first simulated conversation sub-unit, configured to perform a simulated conversation on the target question through at least two agents based on the target conversation turn number indicated by the target parameter, wherein a first agent is configured to generate an answer through a first model, and a second agent is configured to generate a predicted question through a second model, and both the first agent and the second agent belong to the at least two agents;

[0220] An information obtaining sub-unit, configured to obtain the target answer obtained by the first agent and the target predicted question obtained by the second agent in the last round of simulated conversation;

[0221] A reply determining sub-unit, configured to determine reply information based on the target answer and the target predicted question.

[0222] In another possible implementation manner, the reply obtaining unit includes:

[0223] A second simulated conversation sub-unit, configured to perform the simulated conversation for the target conversation turn number on the target question through a first agent and a second agent based on the target conversation turn number indicated by the target parameter, the first agent being configured to generate an answer through a first model, and the second agent being configured to generate a predicted question through a second model;

[0224] A conversation control sub-unit, configured to, for each round of simulated conversation through the first agent and the second agent, obtain the current round answer obtained by the first agent in this round of simulated conversation and the current round predicted question obtained by the second agent in this round of simulated conversation, determine the relevance between the current round answer and the current round predicted question and the target question through a third agent, and if the relevance is less than a set threshold, terminate the simulated conversation between the first agent and the second agent, the third agent determining the relevance through a third model;

[0225] A reply obtaining sub-unit, configured to, in response to the end of the simulated conversation between the first agent and the second agent, use a fourth agent to determine reply information from at least one round of answers and at least one round of questions obtained from the simulated conversation, the fourth agent determining the reply information through a fourth model.

[0226] In another possible implementation manner, the reply determining unit includes:

[0227] A model determination subunit, configured to determine a target model that matches the thinking level depth characterized by the target parameter from models corresponding to multiple different thinking level depths;

[0228] A problem processing subunit, configured to obtain reply information for the target problem by using the target model.

[0229] An embodiment of the present application further provides an electronic device. As Figure 9 shown, it shows a schematic structural diagram of a composition of the electronic device. The electronic device at least includes: an output device 901 and a processor 902.

[0230] Wherein, the processor 902 is configured to be in a working state based on an intelligent program. The intelligent program obtains a target problem and a target parameter characterizing the thinking level depth. The thinking level includes multiple thinking levels with different depths, and the target parameter is used to characterize one of the multiple different depths of thinking levels; the intelligent program provides the target parameter and the target problem to the target model, and the intelligent program has the ability to call the target model; the intelligent program obtains the reply information of the target model for the target problem and outputs the reply information through the output device 901, and the reply information matches the thinking level depth characterized by the target parameter.

[0231] In the present application, the intelligent program can be an artificial intelligence assistant built into the operating system of the electronic device, and the intelligent program can be awakened and started in forms such as voice. Of course, the intelligent program can also be a program capable of realizing human-computer dialogue interaction, without specific limitation.

[0232] In the present application, the output device can have various possible forms. When the types of the output device are different, the specific form of outputting the reply information by the output device will also be different.

[0233] For example, in a possible case of the output device, the output device 901 can be a display screen. On this basis, the reply information can be displayed on the display screen.

[0234] In this possible case, information indicating that the intelligent program is in a working state can also be displayed on the display screen. For example, information for prompting that the intelligent program is in a working state is displayed on the display screen. Another example is that in a display bar with a highlighting display effect, information, images or icons for prompting that the intelligent program is in a working state are output in a scrolling manner. Another example is that when the intelligent program is in a working state, an interaction interface corresponding to the intelligent program is displayed so that the user can know that the intelligent program is in a working state.

[0235] Specifically, when the interactive interface corresponding to the intelligent program is displayed, the interactive interface may include an input area for entering questions, and options corresponding to different thinking level depths may also be displayed in the interactive interface. For details, please refer to the previous Figure 2 and related introductions.

[0236] In another possible case of the output device, the output device may be an audio output device, such as a speaker, etc. On this basis, the reply information may be played through the audio output device.

[0237] Of course, in actual applications, the output device may also include a display screen and an audio output device at the same time. On this basis, while the reply information is displayed, the reply information may be played through the audio.

[0238] Among them, the target model called by the intelligent program may be a model locally deployed on the electronic device. For example, the target model may be stored in the local memory of the electronic device, such as Figure 9 As shown, the electronic device may further include a memory 903, and the memory is at least used to store the data of the target model.

[0239] The target model may also be deployed in the cloud. On this basis, the intelligent program may send a call request to the cloud. The call request carries the target question and target parameters, and the call request is used to request to call the target model. Correspondingly, in response to the call request, the cloud determines the reply information of the target question by using the target model and returns it to the intelligent program.

[0240] Of course, the electronic device may further include an input unit 904 such as a keyboard or a mouse, and more or fewer other components, which are not limited herein.

[0241] An embodiment of the present application also provides a computer program product, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement any information processing method provided by the embodiment of the present application.

[0242] An embodiment of the present application also provides a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device is enabled to implement any information processing method provided by the embodiment of the present application.

[0243] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0244] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0245] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0246] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. An information processing method, comprising: Obtaining a target problem; Obtaining a target parameter characterizing the depth of the thinking level, where the thinking level includes multiple thinking levels with different depths, and the target parameter is used to characterize one of the multiple thinking levels with different depths; Based on the target parameter and the target model, obtaining a response message for the target problem, where the response message matches the depth of the thinking level characterized by the target parameter.

2. The information processing method according to claim 1, where the obtaining of the target parameter characterizing the depth of the thinking level includes: Displaying different depth options for the thinking level; In response to a selection operation, obtaining a target parameter characterizing the depth of the thinking level, where the selection operation is a selection of one of the different depth options for the thinking level.

3. The information processing method according to claim 1, where the obtaining of the target parameter characterizing the depth of the thinking level includes: Based on the target problem, obtaining a target parameter characterizing the depth of the thinking level; where the target problem belongs to one of multiple recommended problems, and the multiple recommended problems are prediction problems generated and displayed in the previous session.

4. The information processing method according to claim 2 or 3, where different depths of the thinking level correspond to different numbers of dialogue turns in the simulated dialogue; The obtaining of the response message for the target problem based on the target parameter and the target model includes: Based on the target number of dialogue turns indicated by the target parameter, performing a simulated dialogue for the target problem through the target model to obtain a response message for the target problem.

5. The information processing method according to claim 4, where the performing of the simulated dialogue for the target problem through the target model based on the target number of dialogue turns indicated by the target parameter includes: Based on the target number of dialogue turns indicated by the target parameter, performing a simulated dialogue for the target problem through at least two agents, where the first agent is used to generate an answer through the first model, and the second agent is used to generate a prediction problem through the second model, and both the first agent and the second agent belong to the at least two agents; Obtaining the target answer obtained by the first agent and the target prediction problem obtained by the second agent in the last round of the simulated dialogue; Determining the response message based on the target answer and the target prediction problem.

6. The information processing method according to claim 4, where the performing of the simulated dialogue for the target problem through the target model based on the target number of dialogue turns indicated by the target parameter includes: Based on the target number of dialogue turns indicated by the target parameter, performing the simulated dialogue for the target problem through the first agent and the second agent for the target number of dialogue turns, where the first agent is used to generate an answer through the first model, and the second agent is used to generate a prediction problem through the second model; Each time a round of simulated conversation is carried out between the first intelligent agent and the second intelligent agent, the current round answer obtained by the first intelligent agent in this round of simulated conversation and the current round predicted question obtained by the second intelligent agent in this round of simulated conversation are obtained. The relevance between the current round answer and the current round predicted question and the target question is determined by a third intelligent agent. If the relevance is less than a set threshold, the simulated conversation between the first intelligent agent and the second intelligent agent is terminated, and the third intelligent agent determines the relevance through a third model; In response to the end of the simulated conversation between the first intelligent agent and the second intelligent agent, a fourth intelligent agent is used to determine a reply message from at least one round of answers and at least one round of questions obtained from the simulated conversation. The fourth intelligent agent determines the reply message through a fourth model.

7. The information processing method according to claim 3, wherein obtaining a target parameter representing the depth of the thinking level based on the target question includes: Determining the historical thinking level depth corresponding to the target question based on the historical thinking level depths adopted by the respective multiple recommended questions generated based on the previous session; Updating a value table using a reinforcement learning algorithm based on the reward score associated with the historical thinking level depth corresponding to the target question and the problem domain corresponding to the target question, where the value table is a relationship table between different problem domains and parameters representing different thinking level depths; Based on the problem domain corresponding to the target question, and in combination with the updated value table, determining a target parameter representing the depth of the thinking level.

8. The information processing method according to claim 7, wherein updating the value table using a reinforcement learning algorithm based on the reward score associated with the historical thinking level depth corresponding to the target question and the problem domain corresponding to the target question includes: Obtaining the user identity characteristics of the user; Updating the value table using a reinforcement learning algorithm based on the reward score associated with the historical thinking level depth corresponding to the target question, the problem domain corresponding to the target question, and the user identity characteristics, where the value table is a relationship table between different problem domains and user attribute characteristics and parameters representing different thinking level depths; The step of determining a target parameter representing the depth of the thinking level based on the problem domain corresponding to the target question and in combination with the updated value table includes: Determining a target parameter representing the depth of the thinking level based on the problem domain corresponding to the target question and the user identity characteristics, and in combination with the updated value table.

9. The information processing method according to claim 1, wherein obtaining a reply message for the target question based on the target parameter and a target model includes: Determining a target model that matches the thinking level depth represented by the target parameter from multiple models corresponding to different thinking level depths; Using the target model to obtain a reply message for the target question.

10. An electronic device, comprising: An output device; A processor, which is used to be in a working state based on an intelligent program. The intelligent program obtains a target problem and a target parameter representing the depth of the thinking level. The thinking level includes multiple thinking levels with different depths, and the target parameter is used to represent one of the multiple thinking levels with different depths. The intelligent program provides the target parameter and the target problem to a target model, and the intelligent program has the ability to call the target model. The intelligent program obtains the reply information of the target model for the target problem and outputs the reply information through the output device, and the reply information matches the depth of the thinking level represented by the target parameter.