Dynamic question recommendation method, device and equipment based on HITS algorithm

By dynamically updating the attributes of asked questions and questions to be recommended through the HITS algorithm, the rigidity of static question lists in complex scenarios is solved, and dynamic adaptability and accuracy of question recommendations are achieved.

CN120744236AActive Publication Date: 2025-10-03NEUSOFT CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510858482.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-03
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Existing technologies have difficulty intelligently and dynamically recommending the next question in scenarios with high uncertainty and complexity, resulting in a static question list that cannot meet the real-time recommendation needs of scenarios such as arbitration.

Method used

Using the HITS algorithm, asked questions are abstracted into hub pages and questions to be recommended are abstracted into authoritative pages. By iteratively updating the hub degree of asked questions and the authority degree of questions to be recommended, dynamic recommendation of questions is achieved.

Benefits of technology

The dynamic nature of question recommendation is achieved, which can better adapt to complex and uncertain scenarios and improve the accuracy and context coherence of question recommendation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744236A_ABST
    Figure CN120744236A_ABST
Patent Text Reader

Abstract

The invention discloses a problem dynamic recommendation method, device and equipment based on an HITS algorithm. The method comprises the steps of obtaining a questioned question set and a to-be-recommended question set; initializing the hub degree of each asked question; initializing the authority of each question to be recommended; for each to-be-recommended question, iteratively updating the authority of the to-be-recommended question based on the initial value of the authority of the to-be-recommended question and the latest value of the hub degree of each questioned question; for each questioned question, iteratively updating the hub degree of the questioned question based on the initial value of the hub degree of the questioned question and the latest value of the authority degree of each question to be recommended; and after iteration is finished, generating a question recommendation result according to the authority of each question to be recommended. Questioned questions are abstracted as hub pages in the HITS algorithm, to-be-recommended questions are abstracted as authoritative pages in the HITS algorithm, and dynamic recommendation of the questions is realized through iterative updating of hub degrees and authoritative degrees.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, and device for dynamically recommending questions based on the HITS algorithm. Background Art

[0002] In scenarios such as project management, market research and analysis, policy formulation and evaluation, community building and management, and arbitration, it's often necessary to carefully review materials and refine the questions to be asked. In arbitration, for example, arbitrators must combine multiple text documents to identify core questions that have substantive value or influence on the decision. To save manpower and improve the efficiency of question extraction, some current technologies can generate a static list of questions based on the text documents, displaying recommended questions. However, in arbitration, for example, the trial process is subject to significant uncertainty, and the specific circumstances of the case are often complex, making static lists of questions difficult to adapt to the characteristics of these scenarios. Therefore, there is a real need to intelligently and dynamically recommend the next question based on the previous question, rather than filtering and sorting a pre-generated list of questions. Unfortunately, a technical solution for intelligently and dynamically recommending the next question based on the previous question is currently unavailable. Summary of the Invention

[0003] Based on the above problems, this application provides a dynamic question recommendation method, device and equipment based on the HITS algorithm, the purpose of which is to intelligently and dynamically recommend the next question based on the previous question, so as to meet the question-asking needs in scenarios with high uncertainty and high complexity.

[0004] The embodiments of this application disclose the following technical solutions:

[0005] In a first aspect, the present application provides a method for dynamic question recommendation based on the HITS algorithm, the method comprising:

[0006] Obtaining a set of asked questions and a set of questions to be recommended; wherein each asked question in the set of asked questions is abstracted as a hub page in the HITS algorithm, and each question to be recommended in the set of questions to be recommended is abstracted as an authoritative page in the HITS algorithm;

[0007] Initialize the hub degree of each question asked based on the time of each question asked and the similarity between multiple known questions; the multiple known questions at least include each question asked and each question to be recommended;

[0008] Initialize the authority of each question to be recommended based on its recommendation score;

[0009] For each question to be recommended, the authority of the question to be recommended is iteratively updated based on the initial value of the authority of the question to be recommended and the latest value of the hub degree of each question that has been asked. For each question that has been asked, the hub degree of the question that has been asked is iteratively updated based on the initial value of the hub degree of the question that has been asked and the latest value of the authority of each question to be recommended.

[0010] After the iteration is completed, question recommendation results are generated based on the authority of each question to be recommended.

[0011] A second aspect of the present application provides a question dynamic recommendation device based on the HITS algorithm, the device comprising:

[0012] A question acquisition module is configured to acquire a set of asked questions and a set of questions to be recommended; wherein each asked question in the set of asked questions is abstracted as a hub page in the HITS algorithm, and each question to be recommended in the set of questions to be recommended is abstracted as an authoritative page in the HITS algorithm;

[0013] A first initialization module is configured to initialize the hubness of each question asked based on the time of each question asked and the similarity between multiple known questions; the multiple known questions include at least each question asked and each question to be recommended;

[0014] The second initialization module is used to initialize the authority of each question to be recommended according to the recommendation score of each question to be recommended;

[0015] A first iteration module is configured to iteratively update the authority of each question to be recommended based on an initial value of the authority of the question to be recommended and the latest value of the hub degree of each question that has been asked;

[0016] A second iteration module is configured to iteratively update the hub degree of each question that has been asked based on the initial value of the hub degree of the question and the latest value of the authority of each question to be recommended;

[0017] The recommendation module is used to generate question recommendation results based on the authority of each question to be recommended after the iteration is completed.

[0018] A third aspect of the present application provides a question dynamic recommendation device based on the HITS algorithm, the device comprising: a memory and a processor;

[0019] The memory is used to store computer programs;

[0020] The processor is configured to run the computer program, and when the computer program is run, the steps of the dynamic question recommendation method based on the HITS algorithm as described in the first aspect are executed.

[0021] Compared with the prior art, this application has the following beneficial effects:

[0022] The present application discloses a method, device and equipment for dynamic question recommendation based on the HITS algorithm. The method obtains a set of asked questions and a set of questions to be recommended; initializes the hub degree of each asked question based on the question time of each asked question and the similarity between multiple known questions; initializes the authority of each question to be recommended based on the recommendation score of each question to be recommended; for each question to be recommended, it iteratively updates the authority of the question to be recommended based on the initial value of the authority of the question to be recommended and the latest value of the hub degree of each asked question; for each asked question, it iteratively updates the hub degree of the asked question based on the initial value of the hub degree of the asked question and the latest value of the authority of each question to be recommended; after the iteration is completed, question recommendation results are generated based on the authority of each question to be recommended. In the technical solution of the present application, the concept of the HITS algorithm is used to abstract the asked questions into the hub page in the HITS algorithm, and the questions to be recommended are abstracted into the authority page in the HITS algorithm. The asked questions have a hub degree, and the questions to be recommended have an authority degree. The hub degree of the asked questions and the authority of the questions to be recommended are interactively iteratively updated. Finally, the latest recommended questions are determined based on the authority of the questions to be recommended after iteration. Since the set of asked questions and the set of questions to be recommended are continuously updated as the questioning process progresses, the dynamic question recommendation method based on the HITS algorithm in this application can match this dynamic change and realize dynamic recommendation of questions, overcoming the shortcomings of rigid and inflexible question recommendation methods, thereby better meeting the questioning needs in scenarios with high uncertainty and high complexity, such as the questioning needs in arbitration scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0024] Figure 1 A schematic diagram of a static list of questions;

[0025] Figure 2 An example effect diagram of implementing dynamic question recommendation using the dynamic question recommendation method based on the HITS algorithm provided in an embodiment of the present application;

[0026] Figure 3 A flowchart of a method for dynamic question recommendation based on the HITS algorithm provided in an embodiment of the present application;

[0027] Figure 4 An example problem association diagram provided in an embodiment of the present application;

[0028] Figure 5 A flowchart for constructing a list of questions to be recommended provided in an embodiment of the present application;

[0029] Figure 6 This is a structural diagram of a question dynamic recommendation device based on the HITS algorithm provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] In scenarios such as project management, market research and analysis, policy formulation and evaluation, community building and management, and arbitration, there are high requirements for the ability and work experience of manually refining questions based on text materials. In order to improve efficiency and save labor costs, it is currently possible to provide a recommended set of questions for manual selection and use based on a specific text material through research on the text material or similar past materials. This type of question set can be presented in the form of a list, and the order in which the questions are presented in the list reflects the degree of recommendation. Often, this type of list-based question set is static, that is, once a specific text material is determined, the list of recommended questions for it is fixed and does not change. Figure 1 A diagram of a static list of questions. Figure 1 In the figure, 8 questions are presented from top to bottom, represented by Q1 to Q8. Figure 1 In the example, Q1 is the most recommended, Q2 is the second, and so on, Q8 is the least recommended. Figure 1 The questions in the question list are automatically recommended in an intelligent manner after analyzing the text materials. Therefore, arbitrators can select questions based on this list and ask them during the court investigation. While the aforementioned automated technology improves the efficiency of question extraction and questioning, and reduces labor costs, its suitability for scenarios with high uncertainty and complexity still needs to be improved.

[0031] Taking arbitration as an example, due to the uncertainty of the trial process and the complexity of the specific circumstances of the case, Figure 1 The static question list shown cannot quickly recommend the next appropriate question to the arbitrator based on the previous question. This makes it difficult for the static question list to meet the arbitrator's real-time recommendation needs. Therefore, it is necessary to propose a technology that can intelligently and dynamically recommend the next question based on the previous question, so as to meet the questioning needs in scenarios with high uncertainty and high complexity.

[0032] In this regard, the inventor proposes in this application to use the concept of the HITS algorithm to abstract the asked questions into the hub page in the HITS algorithm, and to abstract the questions to be recommended into the authoritative page in the HITS algorithm. The asked questions have a hub degree, and the questions to be recommended have an authority degree. The hub degree of the asked questions and the authority degree of the questions to be recommended are interactively updated iteratively. Finally, the latest recommended questions are determined based on the authority degree of the questions to be recommended after the iteration. After each question is asked, the set of asked questions will be updated, and if the new question is from the set of questions to be recommended, the set of questions to be recommended will also be updated after the question is asked. The update of the set of asked questions and the update of the set of questions to be recommended affect the recommendation of the next question. The dynamic question recommendation method based on the HITS algorithm in this application can match this dynamic change and realize the dynamic recommendation of questions, overcoming the shortcomings of the existing question recommendation methods that are rigid and inflexible, thereby better meeting the questioning needs in scenarios with strong uncertainty and high complexity, such as the questioning needs in arbitration scenarios.

[0033] Figure 2 This is an example effect diagram of implementing dynamic question recommendation using the dynamic question recommendation method based on the HITS algorithm provided in the embodiment of this application. Figure 2 In this example, suppose that three moments t1, t2, and t3 need to be used to recommend questions to users based on the set of questions asked and the set of questions to be recommended at the corresponding moments. Among them, t1 is earlier than t2, and t2 is earlier than t3. Figure 2 As shown in , at different moments, as the questioning process progresses, the set of questions asked changes, and the set of questions to be recommended also changes. Figure 2 The Q in the Chinese character “Q” stands for question, and the numbers or characters next to the Q are used to distinguish different questions.

[0034] Figure 2 In this example, at time t1, three questions are recommended from the set of questions to be recommended (including 19 questions from Q1 to Q19), namely Q1, Q3 and Q11. Among them, Q1 is selected as the next question to be asked after Qx7. Therefore, at time t2 after Q1 has been asked, it can be seen that Figure 2 The set of questions asked has been increased by Q1, and accordingly, the set of questions to be recommended no longer has Q1. Based on the set of questions asked and the set of questions to be recommended at time t2, the next question is recommended for Q1 again, and the following is generated: Figure 2 The question recommendation results shown in the figure include Q3, Q4 and Q15. Among them, Q3 is selected as the next question to be asked after Q1. Therefore, at time t3 after Q3 has been asked, the set of questions asked and the set of questions to be recommended change again. At time t3, the following is generated: Figure 2The question recommendation results shown in include Q4, Q12 and Q17. Figure 2 As can be seen in the example, the question recommendation results will change with the new questions asked. In other words, the question recommendation results change dynamically based on the question-asking process. The dynamic properties of this question recommendation mechanism make the recommended questions more closely related to the previous questions, and can better meet the user's question-asking needs in scenarios with high complexity and strong uncertainty, and recommend more accurate questions with stronger contextual coherence to users.

[0035] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0036] See also Figure 3 , which is a flow chart of a method for dynamic question recommendation based on the HITS algorithm provided in an embodiment of the present application. Figure 3 As shown in Figure 1, the dynamic recommendation method based on the HITS algorithm includes:

[0037] S301: Obtain a set of asked questions and a set of questions to be recommended.

[0038] To facilitate understanding, let's use the trial investigation phase in an arbitration scenario as an example. An arbitration tribunal is a temporary organization tasked with resolving disputes between two parties. It can handle a variety of dispute types, including civil and commercial arbitration, labor dispute arbitration (also known as labor arbitration), agricultural contract disputes, and maritime arbitration. During the arbitration trial process, the tribunal must ask questions based on the written materials. This phase, known as the trial investigation, is a critical stage in which arbitrators proactively investigate the facts and legal issues of the case. Its significance goes far beyond simply asking questions; it is a core procedure that directly influences the course of the case and the outcome of the arbitration decision, and is a crucial safeguard for ensuring the fairness and efficiency of the arbitration process. In the past, arbitrators had to thoroughly study and accurately grasp the arbitration application, defense, and evidence submitted by both parties, then analyze and formulate questions. Automatic question recommendation can assist arbitrators in this task of questioning. Arbitrators simply need to select questions from the recommended questions and then ask them.

[0039] As the questioning process develops, questions that have been asked can be continuously collected in actual applications, thereby building a set of questions that have been asked. In addition, the concept of a set of questions to be recommended is also proposed in the technical solution of the present application. The set of questions to be recommended contains multiple questions to be recommended, and these questions to be recommended can be questions worthy of recommendation extracted based on the current text material or past text materials of the same scenario type. In the technical solution of the present application, dynamic question recommendation is performed, that is, question recommendation is performed from the current set of questions to be recommended in combination with the latest questions asked. In the present application, the set of questions that have been asked and the set of questions to be recommended can be regarded as the basic materials for realizing dynamic question recommendation in the embodiments of the present application.

[0040] In this application, the concept of the HITS (Hyperlink-Induced Topic Search) algorithm is used. The HITS algorithm is a ranking algorithm based on the link structure of a web page, which is mainly used to evaluate the authority and hubness of a web page in a search engine. In the HITS algorithm, an authority page (Authority) refers to a web page with high content quality and pointed to by multiple hub pages. For example, an authoritative paper page in an academic field will be cited by a large number of research reviews or course materials. A hub page (Hub) is similar to a "navigation page", which refers to a web page that provides a collection of high-quality links, usually pointing to multiple authoritative pages. For example, an academic navigation website or a course resource list page.

[0041] In the technical solution of the present application, each asked question in the set of asked questions is abstracted as a hub page in the HITS algorithm, and each question to be recommended in the set of questions to be recommended is abstracted as an authoritative page in the HITS algorithm. The purpose is to explore and reflect the connection between asked questions and questions to be recommended in this way. Each hub page, that is, the asked questions in this application, should have a hub degree (or hub value); each authoritative page, that is, the question to be recommended in this application, should have an authority degree (or authority value). The following will introduce the initialization method of the hub degree of asked questions and the initialization method of the authority degree of questions to be recommended through S302 and S303 respectively.

[0042] S302: Initialize the hub degree of each question that has been asked based on the time when each question has been asked and the similarity between multiple known questions.

[0043] This application also introduces the concept of known questions. Both the questions in the set of asked questions and the questions to be recommended in the set of questions to be recommended can be considered known questions. That is, the multiple known questions described in this step include at least: each question in the set of asked questions and each question to be recommended in the set of questions to be recommended. The similarity between these known questions can be represented by a calculated value.

[0044] As an example, the first question and the second question are two different questions among multiple known questions. Among them, the first question can be a question that has been asked or a question to be recommended; the second question can be a question that has been asked or a question to be recommended. The similarity between the first question and the second question can be calculated by the Jaccard similarity calculation method. During the Jaccard similarity calculation process, the first question and the second question are respectively converted into word segmentation sets, and the number of intersection words and the number of union elements of the two word segmentation sets are used to specifically reflect the similarity between the first question and the second question. The specific calculation method is: obtain the word segmentation set of the first question and the word segmentation set of the second question; calculate the number of intersection words of the word segmentation set of the first question and the word segmentation set of the second question, and calculate the number of union words of the word segmentation set of the first question and the word segmentation set of the second question; and use the ratio of the number of intersection words to the number of union words as the similarity between the first question and the second question. The above calculation method can be expressed by formula (1).

[0045]

[0046] In formula (1), Q i Indicates the first question, Q j Indicates the second question, Words(Q i ) represents the word set of the first question, Words(Q j ) represents the word set of the second question. In formula (1), the numerator on the right side of the equal sign represents Words(Q i ) and Words(Q j ) of the intersection of the words, the denominator on the right side of the equal sign represents Words(Q i ) and Words(Q j ) of the union vocabulary. Sim(Q i ,Q j ) represents the Jaccard similarity between the first question and the second question.

[0047] It should be noted that formula (1) is only an example implementation method for calculating the similarity between two questions and does not limit the similarity calculation method. In actual applications, other methods can be used to calculate and represent the similarity between questions, such as converting the questions into vector form and then calculating the similarity between the two questions using the cosine similarity calculation method.

[0048] In this embodiment of the present application, to initialize the hubness of each question, the time of each question is specifically combined with the similarity between multiple known questions, including each question and each question to be recommended. For any two different questions in the multiple known questions, the similarity between the two questions needs to be calculated.

[0049] The following describes the initialization process of the hub degree of the asked questions.

[0050] In this application, a problem association graph is constructed based on multiple known problems and the similarities between each two known problems. Figure 4 An example problem association diagram provided in an embodiment of the present application. Figure 4 In the problem-association graph shown, circles represent nodes, and each node represents a known problem. The lines connecting nodes are called edges. In practice, edges between two nodes in a problem-association graph are bidirectional, with the direction of the edge representing the direction of the problem's transfer. Figure 4 In the figure, only 6 nodes and the edges between them are shown as examples. In actual applications, the number of nodes in the problem association graph is consistent with the number of known problems, and the nodes, node weights, and edge weights between nodes (referred to as edge weights) in the problem association graph can be updated in real time.

[0051] In order to initialize the hub degree of the questions asked, this application uses multiple known questions as nodes in the question association graph, and determines the time decay weight of each question asked as the initial point weight of the corresponding node in the question association graph according to the order of the question asking time. i / t represents the time decay weight of node i, where t represents the total number of questions asked in the set of questions asked, and v i Indicates that the question has been asked i The order of the time of asking questions among the t questions asked, v i The smaller the value, the earlier the question is asked, the smaller the time decay weight corresponding to the question is, and the more severe the time decay is; on the contrary, v i The larger the value, the later the question is asked, the greater the time decay weight corresponding to the question asked, and the weaker the time decay.

[0052] For example, if the set of asked questions contains 10 asked questions, then the question association graph has 10 nodes representing the aforementioned 10 asked questions. These 10 nodes use time-decayed weights as their initial point weights, where t = 10 in this example. If, among these 10 nodes, the question represented by a certain node is the 9th question asked, then the v value of this node is 9, the time-decayed weight of the question represented by this node is 0.9, and the initial point weight of this node is also 0.9. Similarly, if, among these 10 nodes, the question represented by a certain node is the 3rd question asked, then the v value of this node is 3, the time-decayed weight of the question represented by this node is 0.3, and the initial point weight of this node is also 0.3.

[0053] The method of finding the similarity between two questions has been introduced above. In the embodiment of the present application, the similarity between asked questions and asked questions, between asked questions and questions to be recommended, and between questions to be recommended and questions to be recommended can all be quantified and calculated. In order to more accurately explore the connection between questions in the question association graph, it is necessary to combine the similarities between different known questions as a bridge to construct a relatively complete question association graph. In this application, the similarities between multiple known questions are used as the edge weights between the corresponding two nodes in the question association graph.

[0054] Based on the above introduction, for the convenience of understanding, let’s take a target question Q in the set of questions i As an example, we'll introduce how to initialize the hubness of a target question. The target question is one of the questions in the set of questions. In the question association graph, for ease of description and distinction, the node representing the target question is called the target node, denoted by i.

[0055] For the target node in the problem association graph, determine the set of nodes that directly point to the target node in the problem association graph. Further, obtain the problem weight of each node in the node set, determine the edge weight from each node in the node set to the target node, and obtain the out-degree of each node in the node set. The set of nodes that directly point to the target node can be represented as incoming(i), the nodes in this node set incoming(i) can be represented as n, where n∈incoming(i), the out-degree of each node n in incoming(i) is represented as out_degree(n), the problem weight of each node in incoming(i) is represented as PR(n), and the edge weight from node n to node i is represented as graph(n,i). It should be noted that when calculating the problem weight of node i, the problem weight of each node in the node set incoming(i) is available because node n has a pointing relationship with node i, which means that the problem represented by node n has a problem transfer relationship with the problem represented by node i. Therefore, obtaining PR(n) is achievable.

[0056] Multiply the problem weight of the same node in the node set by the edge weight from the node to the target node, and divide it by the out-degree of the node to get the calculation result corresponding to the node; then add the calculation results corresponding to each node in the node set to get the problem weight of the target node. The problem weight PR(Q i ) is:

[0057]

[0058] On the basis of the obtained question weight of the target question, the implementation method of obtaining the initial value of the hub degree of the target question is expressed by formula (3).

[0059] Hub0(Q i )=PR(Q i )×(v i / t); Formula (3)

[0060] As shown in formula (3), the question weight PR(Q i ) and the initial point weight v of the target node i / t multiplied to obtain the initial value of the hub degree of the target question Hub0(Q i ). Combined with the description above, for each question in the set of asked questions, the hub degree of each question can be initialized by using the formed question association graph in the manner described above.

[0061] In S302 , a detailed introduction is given to the initialization method of the hub degree of the asked question. Correspondingly, in S303 , an introduction is also given to the initialization method of the authority degree of the question to be recommended.

[0062] S303: Initialize the authority of each question to be recommended according to the recommendation score of each question to be recommended.

[0063] It should be noted that in the embodiment of the present application, the recommendation score of each question in the set of questions to be recommended can be obtained in advance. The recommendation score is positively correlated with the degree of recommendation, that is, the higher the recommendation score, the higher the degree of recommendation; conversely, the lower the recommendation score, the lower the degree of recommendation.

[0064] Relying on the recommendation scores of each question to be recommended, the minimum score and the maximum score among the recommendation scores of each question to be recommended in the set of questions to be recommended can be determined. On this basis, the difference between the maximum score and the minimum score can be calculated, and the difference between the maximum score and the minimum score obtained is called the first difference. In addition, for each question to be recommended in the set of questions to be recommended, its difference with the above-mentioned minimum score can also be obtained. In this application, the difference between the recommendation score of each question to be recommended and the minimum score among the recommendation scores of each question to be recommended is called the second difference. The initial value of the authority of each question to be recommended can be calculated by formula (4).

[0065]

[0066] In formula (4), max(scores) and min(scores) represent the maximum score and minimum score of each question to be recommended, respectively. The denominator on the right side of the equal sign in formula (4) represents the first difference. c , and its recommendation score is expressed as score c , then the numerator on the right side of the equation (4) represents the value of the problem to be recommended Q c The second difference between the recommended score and the minimum score is calculated. The ratio of the second difference to the first difference is Auth0(Q c ) as the recommended question Q c The initial value of the authority of each question to be recommended can be calculated based on formula (4).

[0067] In formula (4), the recommendation scores of the recommended questions are actually scaled to the minimum and maximum values. In this way, the recommendation scores of each recommended question are scaled to the interval [0, 1] as the initial value of the authority.

[0068] It should be noted that, in the embodiment of the present application, S302 and S303 may be as follows: Figure 3 The execution order shown in the figure can also be that S303 is executed first and then S302. In addition, S302 and S303 can also be executed simultaneously. Therefore, the execution order of these two steps is not limited in this application.

[0069] Through the execution of S302 and S303, the initialization of the hub degree of each asked question in the set of asked questions and the initialization of the authority degree of each question to be recommended in the set of recommended questions are completed. In an embodiment of the present application, in order to be able to dynamically recommend the next recommended question based on the previously asked question, the hub degree of each asked question and the authority degree of each question to be recommended are further iteratively updated by executing S304. Based on the concept of the HITS algorithm, in the process of iteratively updating the hub degree and authority degree, the update of the hub degree of the asked question depends on the numerical value of the authority degree of each question to be recommended. In addition, the update of the authority degree of the question to be recommended depends on the numerical value of the hub degree of each asked question. The following is a detailed explanation in conjunction with S304.

[0070] S304. For each question to be recommended, the authority of the question to be recommended is iteratively updated based on the initial value of the authority of the question to be recommended and the latest value of the hub degree of each question that has been asked; for each question that has been asked, the hub degree of the question that has been asked is iteratively updated based on the initial value of the hub degree of the question that has been asked and the latest value of the authority of each question to be recommended.

[0071] S304 mainly involves two levels of iterative updates: one level is for updating the authority of questions to be recommended, and the other level is for updating the hubness of questions that have been asked. These are described below.

[0072] First, we introduce the update of the authority of the question to be recommended. Taking the target question to be recommended as an example, we introduce the technical implementation of iteratively updating the authority of the target question to be recommended based on the initial value of the authority of the target question to be recommended and the latest value of the hub degree of each question that has been asked. The target question to be recommended is one of the questions to be recommended in the set of questions to be recommended, with Q c To express.

[0073] For the target recommendation problem Q c , obtain the transition probability W of each question asked to the target question to be recommended i→c In this transition probability, the subscript i→c indicates that the transition from question Q i Move to question Q c .Q i Represents any question in the set Asked. It's important to note that the transition probability between questions indicates the likelihood of a logical connection between the previous and the next question. A larger transition probability indicates a higher likelihood of a logical connection; conversely, a smaller transition probability indicates a lower likelihood of a logical connection. Optional methods for calculating transition probabilities will be discussed later in this article and are not detailed here.

[0074] Formula (5) shows the iterative update target recommendation problem Q c How to achieve authority:

[0075]

[0076] In formula (5), Hub(Q i ) (k) Represents the updated question Q after the kth iteration i . α is the preset weight coefficient, and its value can be set according to actual needs. c ) (k+1) It means the target recommendation question Q updated after the k+1th iteration c Authoritativeness. c ) is the target recommendation question Q c The initial value of the authority. It is not difficult to see from formula (5) that the target recommendation question Q c Each iteration of the authority needs to depend on the target recommendation question Q c The initial value of the authority, and the dependence of each question Q in the set of questions Askedi The hub degree updated after the previous iteration, and each question Q i The transition probability W to the target recommendation problem i→c .

[0077] The implementation process of formula (5) is briefly described in text below. In order to iteratively update the authority of the target question to be recommended, it is necessary to obtain the latest value of the hub degree of each question asked and the first product of the transition probability from the corresponding question asked to the target question to be recommended. The first product corresponding to each question asked is summed to obtain the first sum calculation result

[0078] Get the first weight coefficient α and the first sum calculation result The second product of , that is, the first factor on the right side of the equal sign of formula (5) is obtained. The second weight coefficient (1-α) and the initial value Auth0(Q c ) is the third product of the equation (5), which gives the second factor on the right side of the equal sign. The sum of the first weight coefficient α and the second weight coefficient (1-α) is 1. Finally, the second product and the third product are summed to obtain the second summation result Auth(Q c ) (k+1) , and the second summation result Auth(Q c ) (k+1) The authority of the target recommendation question after the current iteration update.

[0079] Next, we will continue to introduce the update of the hub degree of the question that has been asked. Taking the target question as an example, we will introduce the technical implementation of iteratively updating the hub degree of the target question based on the initial value of the hub degree of the target question and the latest value of the authority of each question to be recommended. The target question is one of the questions in the set of questions that have been asked, with Q as the number. i To express.

[0080] Question Q has been asked about the target i , obtain the transition probability W of the target question to each question to be recommended i→c In this transition probability, the subscript i→c indicates that the transition from question Q i Move to question Q c .Q c Represents any question to be recommended in the set of Candidates.

[0081] Formula (6) shows the iterative update target question Q i The implementation method of hub degree is:

[0082]

[0083] In formula (6), Auth(Q c ) (k) Represents the updated recommendation question Q after the kth iteration c The authority of Hub(Q i ) (k+1) It means that the updated target question Q has been asked after the k+1th iteration i Hub degree. Hub0(Q i ) is the target question Q i The initial value of the hub degree. It is not difficult to see from formula (6) that the target has asked question Q i Each iteration of the hub degree needs to depend on the target question Q i The initial value of the hub degree, and the dependence of each recommended question Q in the set of recommended questions Candidates c The authority updated after the previous iteration, and the target question Q i The transition probability W to each recommended question i→c .

[0084] The implementation process of formula (6) is briefly described in text below. In order to iteratively update the hub degree of the target question, it is necessary to obtain the latest value of the authority of each question to be recommended and the fourth product of the transition probability from the target question to the corresponding question to be recommended. The fourth product corresponding to each question to be recommended is summed to obtain the third summation calculation result

[0085] Get the first weight coefficient α and the third summation calculation result The fifth product of , we get the first factor on the right side of the equation (6). Get the initial value Hub0(Q i ) is the sixth product, which gives the second factor on the right side of the equal sign in formula (6). The sum of the first weight coefficient α and the second weight coefficient (1-α) is 1. Finally, the sum of the fifth product and the sixth product gives the fourth summation result Hub(Q i ) (k+1) and the fourth summation result Hub(Q i ) (k+1) As the hub degree of the target question after the iteration update.

[0086] S305: After the iteration is completed, a question recommendation result is generated according to the authority of each question to be recommended.

[0087] Combined with the description of the iterative update of the hub degree of the asked questions and the iterative update of the authority of the questions to be recommended in S304 above, the specific calculation method can be known. The iterative update meets the needs of interactive optimization of the hub degree of the asked questions and the authority of the questions to be recommended. In actual applications, the number of iterations can be set according to the optimization demands or the real-time requirements of the dynamic recommendation questions. As an example, the iteration can be set to end after 5 times. In this example, after completing 5 iterations, the authority of each question to be recommended in the set of questions to be recommended can be obtained. Based on this, the question recommendation results can be generated according to the authority of each question to be recommended at this time.

[0088] For example, after the iteration is complete, the top X recommended questions with the highest authority and their corresponding authority scores are included in the question recommendation results, and the question recommendation results are output. X is a positive integer, and X is less than the total number of recommended questions in the set of recommended questions. For example, if X = 3, the top three recommended questions with the highest authority scores are included in the question recommendation results.

[0089] For another example, after the iteration is complete, the top Y% of recommended questions with the highest authority and their corresponding authority are included in the question recommendation results, and the question recommendation results are output. Y is a positive number less than 100. For example, Y = 5, which means that the top 5% of the recommended questions with the highest authority are recommended.

[0090] By displaying authority, we can reflect the numerical differences in the authority of recommended questions, making it easier for users to reference and select questions for their next question. Of course, it is also possible to not display authority. For example, in the question recommendation list, the recommended questions can be sorted in descending order of authority, so that users can also understand the relative authority of the questions.

[0091] In practical applications, to avoid recommending duplicate questions and improve the efficiency and effectiveness of dynamic question recommendation, after generating a question recommendation result based on the authority of each question to be recommended, the method can further record the next question selected in the question recommendation result; then, the next question is removed from the set of questions to be recommended and added to the set of questions already asked. This completes the update of both the set of questions already asked and the set of questions to be recommended.

[0092] In the technical solution of the present application, the concept of the HITS algorithm is used to abstract the asked questions into the hub page in the HITS algorithm, and the questions to be recommended are abstracted into the authoritative page in the HITS algorithm. The asked questions have a hub degree, and the questions to be recommended have an authority degree. The hub degree of the asked questions and the authority degree of the questions to be recommended are interactively and iteratively updated. Finally, the latest recommended questions are determined based on the authority degree of the questions to be recommended after the iteration. Since the set of asked questions and the set of questions to be recommended are continuously updated as the questioning process progresses, the dynamic question recommendation method based on the HITS algorithm of the present application can match this dynamic change and realize dynamic recommendation of questions, overcoming the drawbacks of rigid and inflexible question recommendation methods, thereby better meeting the questioning needs in scenarios with strong uncertainty and high complexity, such as the questioning needs in arbitration scenarios.

[0093] Regarding the set of questions to be recommended mentioned in the above embodiment, the following describes a method for constructing a set of questions to be recommended (a list of questions to be recommended) in the form of a list. In the embodiment of this application, each question to be recommended in the list of questions to be recommended is arranged in descending order according to the recommendation score. The form of the list can be referred to in Figure 1 Essentially, for a specific text material, the list of questions to be recommended is static and will not change, but it can assist in determining the initial value of the authority of each question to be recommended in the technical solution of this application.

[0094] Figure 5 This is a flowchart of a method for constructing a list of questions to be recommended provided in an embodiment of the present application. Figure 5 As shown in Figure 2, building a list of questions to be recommended includes the following steps:

[0095] S501 , using the keywords and word frequencies of the target text material, retrieve multiple feature vectors from a pre-constructed co-occurrence matrix.

[0096] In the embodiments of the present application, questions and recommendations are all based on the target text material. Taking the trial investigation link in the arbitration scenario as an example, the type of target text material can specifically refer to the trial transcript. With respect to the target text material, keywords and the frequency of each keyword in the target text material can be extracted. A pre-built keyword library can be used to extract keywords. For example, the words in the target text material that hit the keyword library are determined as keywords. The frequency of a keyword can be calculated by taking the ratio of the number of times a keyword appears in the target text material to the sum of the number of times all keywords appear in the target text material as the frequency of the keyword.

[0097] In the embodiments of this application, a co-occurrence matrix is ​​pre-constructed based on a collection of historical text materials. The co-occurrence matrix is ​​a two-dimensional matrix, where the first dimension represents questions and the second dimension represents keywords. The questions represented by the first dimension of the co-occurrence matrix are question statements extracted from the collection of historical text materials, and the keywords represented by the second dimension of the co-occurrence matrix are text keywords extracted from the collection of historical text materials. The first and second dimensions are used to distinguish the two different dimensions of the co-occurrence matrix: rows and columns. For example, if the first dimension is the rows of the matrix and the second dimension is the columns of the matrix, then each row of the co-occurrence matrix corresponds to a question, and each column corresponds to a keyword. In another example, if the first dimension is the columns of the matrix and the second dimension is the rows of the matrix, then each column of the co-occurrence matrix corresponds to a question, and each row corresponds to a keyword. For easier understanding and description, the following embodiments will be described using an example where the first dimension is the rows of the matrix and the second dimension is the columns of the matrix. It should be understood that the technical concepts described in this application's technical solution can still be applied adaptively even when the first dimension is the columns of the matrix and the second dimension is the rows of the matrix.

[0098] The co-occurrence matrix is ​​established on the basis of a large amount of historical text materials (of the same type as the target text material) in the historical text material set by analyzing the co-occurrence of keywords and questions, and reflects the closeness of the connection between keywords and questions in the historical text materials in the form of a matrix. The larger the value of the element at a certain position in the co-occurrence matrix, the closer the connection between the corresponding keyword and the question. In this application, the co-occurrence matrix is ​​a data set that reflects the closeness of the connection between keywords and questions based on historical data. It can be used to recommend questions suitable for asking questions for new text materials (i.e., target text materials), and then form a list of questions to be recommended for the target text material.

[0099] Based on the constructed co-occurrence matrix, the keywords in the target text material can be used to complete data indexing in the co-occurrence matrix, obtaining multiple feature vectors that correspond one-to-one with the keywords in the target text material. As an example, if the five keywords in the target text material hit the keyword library, five feature vectors can be extracted from the co-occurrence matrix accordingly. In the example where the first dimension is the rows of the matrix and the second dimension is the columns of the matrix, the values ​​of the elements at each position in the column corresponding to a keyword together constitute a feature vector corresponding to that keyword.

[0100] S502: Perform feature enhancement processing on feature vectors corresponding to corresponding keywords using word frequencies to obtain multiple enhanced feature vectors.

[0101] In the specific implementation of this step, feature vectors can be enhanced by word frequency (for example, by multiplying the feature vector corresponding to a keyword by the keyword's word frequency in the target text material). This strengthens the critical role of word frequency, thereby creating a strong correlation between the enhanced feature vector and the inherent characteristics of the target text material. The enhanced feature vector not only reflects the co-occurrence characteristics of questions and keywords in the historical text material, but also reflects the important contribution of keywords in the target text material to question extraction and recommendation.

[0102] S503: Perform a sum operation on the multiple enhanced feature vectors to obtain a result vector.

[0103] The length of each enhanced feature vector is consistent. In this step, the values ​​of the elements at the same position of these vectors are accumulated to obtain a total vector, which is the result vector.

[0104] S504: Generate a list of questions to be recommended for the target text material according to the recommendation scores in the result vector.

[0105] Each element in the result vector has a one-to-one correspondence with a question, and the value of the element at each position can be used to represent the recommended score for the corresponding question. Questions corresponding to positions with larger values ​​in the result vector can be selected and sorted in descending order of recommended score to generate a list of recommended questions for the target text material.

[0106] For example, the resulting vector from the summation operation is represented as [11 13 25 0 0 10], where each position corresponds to question 1, question 2, question 3, question 4, question 5, and question 6. The list of recommended questions that can be formed from this resulting vector is, in descending order of recommendation score, as follows: ① Question 3; ② Question 2; ③ Question 1; ④ Question 6.

[0107] It should be noted that the technical solutions introduced above involve the processing or conversion of questions and question-related data many times. In practical applications, considering that there may be semantically similar connections between questions with different expressions, if questions with different expressions are treated separately, the dynamic recommendation efficiency of questions may be reduced. To this end, questions can be standardized in advance. For example, a similar question library can be constructed, from which representative questions and similar questions with a high degree of semantic similarity to the representative questions can be queried or indexed. When recommending questions, only representative questions can be considered, thereby standardizing the related processing of various similar questions into the processing of representative questions.

[0108] In step S304 described in the previous embodiment, the iterative updating process for the authority of the question to be recommended and the hubness of the question already asked was detailed. Combining formulas (5) and (6), it can be seen that these updates rely on the calculation results of the transition probabilities between questions. For ease of understanding, the following describes how to obtain the transition probabilities.

[0109] The calculation of transition probabilities relies on the construction of the question transition matrix. In this application, the transition probabilities between corresponding questions are extracted from the question transition matrix. Let's first introduce the construction method of question transition probabilities:

[0110] First, the historical text material collection is traversed, and the number of adjacent questions in each historical text material that are asked in sequence is counted. The count value is added to the corresponding position of the adjacent questions in the blank first transfer matrix until the traversal is completed, thereby obtaining the second transfer matrix. The first dimension and the second dimension of the first transfer matrix both represent questions. It is called the blank first transfer matrix because the first transfer matrix is ​​considered an initialized question transfer matrix. In the first transfer matrix, the elements at each position are set to 0. In one example, the first dimension and the second dimension respectively represent the rows and columns of the matrix. For example, the so-called adjacent questions that are asked in sequence in a historical text material are: question i, question j, and question k appear in sequence in a certain historical text material, where question i and question j are adjacent questions one after the other, and question j and question k are adjacent questions one after the other. Based on this, in the first transfer matrix, the position at the i-th row and j-th column is counted once, and the position at the j-th row and k-th column is counted once. When determining the position to be counted, the preceding question is indexed according to the first dimension (rows), and the following question is indexed according to the second dimension (columns), thereby determining the counting position in the first transfer matrix. By traversing the historical text materials and counting, a second transfer matrix is ​​formed. The numerical values ​​in the second transfer matrix reflect the close connection between the contexts of the questions. The closeness of the contextual connection between the questions is positively correlated with the numerical values ​​in the second transfer matrix.

[0111] The second transfer matrix is ​​numerically normalized according to the first dimension to obtain the problem transfer matrix. Through normalization, the sum of the transition probabilities of each problem in the first dimension in the final problem transfer matrix is ​​1. That is, in the problem transfer matrix, the sum of the numerical values ​​of the elements at each position in a row is 1, and the sum of the transition probabilities from the problem corresponding to the row to the other problems represented by the columns of the problem transfer matrix is ​​1. Taking the transition probability from the first problem to the second problem as an example, it is only necessary to extract the numerical value of the element corresponding to the first problem in the first dimension and the second problem in the second dimension in the problem transfer matrix. This numerical value is the transition probability from the first problem to the second problem.

[0112] The above method of constructing a question transfer matrix and extracting the transition probability from one question to another is applicable to scenarios where the question exists in the above question transfer matrix. However, in actual applications, it is also possible that the newly proposed question is not a question extracted from the historical question material library. This will result in the problem transfer matrix failing to cover the newly proposed question in the first and second dimensions, that is, it is impossible to extract the transition probability of this question to other questions from the problem transfer matrix. For example, during the trial, the arbitrator may abandon the traditional questions and ask one or more personalized questions based on the details of the case. However, the problem transfer matrix does not have the transition probability from this question to each recommended question, and HITS iteration is even more impossible.

[0113] To address the difficulty in extracting the transition probability from a newly asked question to other questions using the question transition matrix, this application further introduces possible methods for calculating transition probabilities. See the following description for details. The following description uses the example of extracting the transition probability from a newly asked question to another question to be recommended (for ease of explanation, referred to as the target question to be recommended).

[0114] If the question with the latest asking time in the set of asked questions (denoted as Q new ) is not a question extracted from the historical text material library, then the latest question asked to the target to be recommended question (denoted as Q q ) is calculated as:

[0115] Calculate the latest question Q new and other asked questions in the asked questions collection r The similarity of Q r Represents the set of questions asked except Q new Any other questions asked except . The method of calculating the similarity between questions has been introduced in the previous article and will not be repeated here. Then, other questions Q r Each to the target to be recommended question Q q The transition probability W r→q Here, it is assumed that every other question Q r Each to the target to be recommended question Q q The transition probability W r→qAll of these can be extracted from the problem transition matrix. Alternatively, this can be understood as follows: once a new problem (not originally included in the problem transition matrix) has its problem transition probability calculated using the method described here to reach other known problems, the existing problem transition matrix can be updated accordingly, for example, the row corresponding to the new problem in the problem transition matrix. At this point, the new problem can also be included in the scope of known problems.

[0116] Next, the question Q with the latest question time will be new and other asked questions in the asked questions collection r Similarity to other questions Q r Each to the target to be recommended question Q q The transition probability W r→q , and the attenuation coefficient, according to other questions Q r The corresponding multiplication and summation are performed to obtain the fifth summation calculation result as the preliminary transition probability from the question asked at the latest question time to the target question to be recommended, which can be seen in formula (7):

[0117]

[0118] In formula (7), λ represents the attenuation coefficient. Multiplying with the attenuation coefficient is mainly due to the fact that the transfer probability W in this application new→q It is through W r→q Obtained indirectly by calculation, this indirect calculation process can be regarded as a certain degree of attenuation compared to the transfer probability obtained directly. As an example, the attenuation coefficient λ = 0.5, and the specific value of the attenuation coefficient here can be set by yourself and is not limited. The above method of calculating the transfer probability between problems uses the SimRank core mechanism of intermediate node propagation similarity, assuming that similar nodes have similar neighbors. In this application, this mechanism can also be understood as similar problems pointing to similar problems. In addition, the attenuation coefficient is used to control the intensity of propagation.

[0119] However, the fifth summation result obtained by calculation in the current formula (7) can represent the problem Q new To the target recommendation question Q q The transfer tendency of , but the calculated result is not in the range of [0,1]. This is because it has not undergone normalization processing and therefore cannot represent the true transfer probability. For this reason, it is also necessary to new To each problem to be recommended (for example, the target problem to be recommended Q q )' new→q Perform normalization operation, specifically: first, Q newThe initial transition probabilities of each question to be recommended are summed to obtain the result SUM (i.e., the sum of the initial transition probabilities), see formula (8); then, the result SUM obtained by summing formula (8) is used as the denominator, and the result calculated by formula (7) is used as the numerator, thereby normalizing the result calculated by formula (7) to obtain the problem Q new To the target recommendation question Q q The normalized transition probability W new→q , see formula (9):

[0120]

[0121] Using formula (8), considering the same problem Q new The overall situation of the transfer tendency to each recommended question. The normalization of the initial transfer probability is completed using formula (9). The normalized result is in the interval [0,1] and can be used to represent the transfer probability. According to the above formula, even if there is no question in the question transfer matrix (for example, question Q new ) to each question to be recommended, we can also combine the transition probabilities of similar questions in the set of questions to be recommended to indirectly obtain the question Q new The transition probability to each recommended question.

[0122] This application utilizes the HITS algorithm in combination with the transition probability and a static question recommendation list, and captures the implicit logical chain between questions in real time through the bidirectional iteration of the hub degree of the asked questions and the authority of the questions to be recommended. And even if there is no direct edge in the question transition matrix for some question pairs, and the transition probability cannot be directly read, based on the multi-order propagation of the corresponding nodes of the asked questions in the question association graph, it is still possible to establish an association and indirectly calculate the transition probability, thereby realizing dynamic recommendation of questions. Therefore, when the user suddenly switches topics (such as from "labor contract" to "work injury insurance"), the hub degree of the asked questions can be quickly redistributed. This dynamic mechanism can achieve a more flexible, dynamic, and adaptable question recommendation effect than relying solely on the static question transition matrix.

[0123] Based on the HITS algorithm-based question dynamic recommendation method described in the above embodiment, the present application also provides a HITS algorithm-based question dynamic recommendation device. The implementation of the device will be described below with reference to the accompanying drawings. Figure 6 This is a schematic diagram of the structure of a dynamic question recommendation device based on the HITS algorithm provided in an embodiment of the present application. Figure 6 As shown, the dynamic recommendation device based on the HITS algorithm includes:

[0124] The question acquisition module 601 is used to acquire a set of asked questions and a set of questions to be recommended; wherein each asked question in the set of asked questions is abstracted as a hub page in the HITS algorithm, and each question to be recommended in the set of questions to be recommended is abstracted as an authoritative page in the HITS algorithm;

[0125] A first initialization module 602 is configured to initialize the hubness of each question asked based on the time of each question asked and the similarity between multiple known questions; the multiple known questions include at least each question asked and each question to be recommended;

[0126] The second initialization module 603 is used to initialize the authority of each question to be recommended according to the recommendation score of each question to be recommended;

[0127] A first iterative module 604 is configured to iteratively update the authority of each question to be recommended based on the initial value of the authority of the question to be recommended and the latest value of the hub degree of each question that has been asked;

[0128] The second iterative module 605 is configured to iteratively update the hub degree of each question that has been asked based on the initial value of the hub degree of the question and the latest value of the authority of each question to be recommended;

[0129] The recommendation module 606 is used to generate question recommendation results according to the authority of each question to be recommended after the iteration is completed.

[0130] In an optional implementation, the first iteration module 604 includes:

[0131] A first transition probability acquisition unit is configured to acquire, for a target question to be recommended, a transition probability from each question already asked to the target question to be recommended; the target question to be recommended being one of the questions to be recommended in the set of questions to be recommended;

[0132] A first calculation unit is used to obtain the latest value of the hub degree of each asked question and the first product of the transition probability of the corresponding asked question to the target question to be recommended;

[0133] a second calculation unit, configured to sum the first products corresponding to the questions asked to obtain a first sum calculation result;

[0134] a third calculation unit, configured to obtain a second product of the first weight coefficient and the first sum calculation result, and obtain a third product of the second weight coefficient and the initial value of the authority of the target question to be recommended; the sum of the first weight coefficient and the second weight coefficient is 1;

[0135] The fourth calculation unit is used to sum the second product and the third product to obtain a second sum calculation result, and use the second sum calculation result as the authority of the target question to be recommended after the iterative update.

[0136] In an optional implementation, the second iteration module 605 includes:

[0137] a second transition probability acquisition unit, configured to acquire, for a target asked question, a transition probability from the target asked question to each question to be recommended; the target asked question being one of the asked questions in the set of asked questions;

[0138] A fifth calculation unit, configured to obtain a fourth product of the latest value of the authority of each question to be recommended and a transition probability from the target asked question to the corresponding question to be recommended;

[0139] a sixth calculation unit, configured to sum the fourth products corresponding to the questions to be recommended to obtain a third sum calculation result;

[0140] a seventh calculation unit, configured to obtain a fifth product of the first weight coefficient and the third sum calculation result, and obtain a sixth product of the second weight coefficient and the initial value of the hub degree of the target asked question; the sum of the first weight coefficient and the second weight coefficient is 1;

[0141] An eighth calculation unit is configured to sum the fifth product and the sixth product to obtain a fourth sum calculation result, and use the fourth sum calculation result as the hub degree of the target asked question after the iterative update.

[0142] In an optional implementation, the first transition probability acquisition unit or the second transition probability acquisition unit acquires the transition probability between two questions in the following manner:

[0143] Traversing the historical text material set, counting the number of adjacent questions in each historical text material that are asked in sequence, and adding the count value to the corresponding position of the two adjacent questions in the first transfer matrix until the traversal is completed, thereby obtaining a second transfer matrix; the first dimension and the second dimension of the first transfer matrix both represent questions, and the elements at each position are set to 0. The first transfer matrix is ​​an initialization structure of the question transfer matrix;

[0144] Normalizing the second transfer matrix according to the first dimension to obtain a problem transfer matrix;

[0145] If it is necessary to obtain the transition probability from the first question to the second question, extract the value of the element corresponding to the first question in the first dimension and the second question in the second dimension in the question transfer matrix as the transition probability from the first question to the second question.

[0146] In an optional implementation, if the latest question in the set of asked questions is not a question extracted from the historical text material library, the transition probability from the latest question to the target question to be recommended is calculated as follows:

[0147] Calculating the similarity between the question asked the latest and other questions asked in the set of questions;

[0148] Extracting the transition probabilities of other asked questions to the target question to be recommended from the question transition matrix;

[0149] Multiply the similarities between the latest asked question and other asked questions in the set of asked questions, the transition probabilities of each of the other asked questions to the target question to be recommended, and the attenuation coefficients according to the corresponding values ​​of the other asked questions, and sum them up to obtain a fifth summation result as the preliminary transition probability from the latest asked question to the target question to be recommended;

[0150] The fifth summation calculation result is divided by the sum of the preliminary transition probabilities for normalization, and the calculated quotient is used as the normalized transition probability from the question asked at the latest question time to the target question to be recommended; the sum of the preliminary transition probabilities is the sum of the preliminary transition probabilities from the question asked at the latest question time to each question to be recommended.

[0151] In an optional implementation, the first initialization module 602 includes:

[0152] a node initial point weight determination unit, configured to use the plurality of known questions as nodes in a question association graph, and determine, based on the order of questioning time of each question, a time-decay weight of each question as the initial point weight of the corresponding node in the question association graph;

[0153] an inter-node edge weight determination unit, configured to use the similarities between any two of the plurality of known questions as edge weights between corresponding two nodes in the question association graph;

[0154] The first initialization module 602 is specifically configured to initialize the hub degree of the target asked question in the following manner:

[0155] For a target node in the question association graph, determine a set of nodes that directly point to the target node in the question association graph, further obtain a question weight for each node in the node set, determine an edge weight from each node in the node set to the target node, and obtain an out-degree for each node in the node set; the question corresponding to the target node is the target asked question; and the target asked question is one of the asked questions in the set of asked questions;

[0156] Multiply the problem weight of the same node in the node set by the edge weight from the node to the target node, and divide the result by the out-degree of the node to obtain the calculation result corresponding to the node;

[0157] Adding the calculation results corresponding to each node in the node set to obtain the problem weight of the target node;

[0158] The question weight of the target node is multiplied by the initial point weight of the target node to obtain the initial value of the hub degree of the target asked question.

[0159] In an optional implementation, the second initialization module 603 includes:

[0160] A score determination unit, used to determine the minimum score and the maximum score among the recommended scores of each question to be recommended;

[0161] a difference calculation unit, configured to obtain a first difference between the maximum score and the minimum score;

[0162] a difference obtaining unit, configured to obtain, for each question to be recommended, a second difference between the recommendation score of the question to be recommended and the minimum score;

[0163] An initial value determining unit is configured to use a ratio of the second difference to the first difference as an initial value of the authority of the question to be recommended.

[0164] In an optional implementation, the similarity between the first question and the second question is calculated as follows:

[0165] Obtain a word segmentation set for the first question and a word segmentation set for the second question;

[0166] Calculate the number of words in the intersection of the word set of the first question and the word set of the second question, and calculate the number of words in the union of the word set of the first question and the word set of the second question;

[0167] The ratio of the number of the intersection words to the number of the union words is used as the similarity between the first question and the second question.

[0168] In an optional implementation, the set of questions to be recommended is in the form of a list of questions to be recommended, in which the questions to be recommended are arranged in descending order of recommendation scores. The dynamic question recommendation device based on the HITS algorithm further includes a question list construction module for constructing the list of questions to be recommended in the following manner:

[0169] Using the keywords and word frequencies of the target text material, a plurality of feature vectors are retrieved from a pre-constructed co-occurrence matrix; the feature vectors correspond one-to-one to the keywords of the target text material; the co-occurrence matrix is ​​a two-dimensional matrix, wherein the first dimension represents questions and the second dimension represents keywords; wherein the questions represented by the first dimension of the co-occurrence matrix are question sentences extracted from a collection of historical text materials, and the keywords represented by the second dimension of the co-occurrence matrix are text keywords extracted from the collection of historical text materials;

[0170] Performing feature enhancement processing on feature vectors corresponding to corresponding keywords using the word frequencies to obtain multiple enhanced feature vectors;

[0171] Performing a sum operation on the multiple enhanced feature vectors to obtain a result vector;

[0172] A list of to-be-recommended questions for the target text material is generated according to the recommendation scores in the result vector.

[0173] In an optional implementation, the recommendation module 606 is specifically configured to:

[0174] After the iteration is completed, the top X recommended questions with higher authority and their corresponding authority are included in the question recommendation result, and the question recommendation result is output; or after the iteration is completed, the top Y% recommended questions with higher authority and their corresponding authority are included in the question recommendation result, and the question recommendation result is output;

[0175] X is a positive integer, and X is less than the total number of questions to be recommended in the set of questions to be recommended; Y is a positive number less than 100.

[0176] In an optional implementation, the question dynamic recommendation device based on the HITS algorithm further includes a set updating module for:

[0177] Recording the next question selected in the question recommendation result;

[0178] The next question is deleted from the set of questions to be recommended, and the next question is included in the set of questions that have been asked.

[0179] Based on the methods and devices described in the above embodiments, this application further provides a dynamic question recommendation device based on the HITS algorithm. The device includes: a memory and a processor;

[0180] The memory is used to store computer programs;

[0181] The processor is configured to run the computer program, and when the computer program is run, the steps of the dynamic question recommendation method based on the HITS algorithm as described in any implementation manner in the method embodiment are executed.

[0182] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The device and equipment embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components indicated as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0183] The above is merely one specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A dynamic question recommendation method based on the HITS algorithm, characterized in that: include: Obtaining a set of asked questions and a set of questions to be recommended; wherein each asked question in the set of asked questions is abstracted as a hub page in the HITS algorithm, and each question to be recommended in the set of questions to be recommended is abstracted as an authoritative page in the HITS algorithm; Initialize the hub degree of each question asked based on the time of each question asked and the similarity between multiple known questions; the multiple known questions at least include each question asked and each question to be recommended; Initialize the authority of each question to be recommended based on its recommendation score; For each question to be recommended, the authority of the question to be recommended is iteratively updated based on the initial value of the authority of the question to be recommended and the latest value of the hub degree of each question that has been asked. For each question that has been asked, the hub degree of the question that has been asked is iteratively updated based on the initial value of the hub degree of the question that has been asked and the latest value of the authority of each question to be recommended. After the iteration is completed, question recommendation results are generated based on the authority of each question to be recommended.

2. The method according to claim 1, characterized in that For each question to be recommended, the authority of the question to be recommended is iteratively updated based on the initial value of the authority of the question to be recommended and the latest value of the hub degree of each question that has been asked, including: For the target question to be recommended, obtaining the transition probability of each asked question to the target question to be recommended; the target question to be recommended is one of the questions to be recommended in the set of questions to be recommended; Obtain the latest value of the hub degree of each asked question and the first product of the transition probability from the corresponding asked question to the target question to be recommended; Summing the first products corresponding to the questions asked to obtain a first sum calculation result; Obtaining a second product of the first weight coefficient and the first sum calculation result, and obtaining a third product of the second weight coefficient and the initial value of the authority of the target question to be recommended; the sum of the first weight coefficient and the second weight coefficient is 1; The second product and the third product are summed to obtain a second sum calculation result, and the second sum calculation result is used as the authority of the target question to be recommended after the iterative update.

3. The method according to claim 1, characterized in that For each question that has been asked, iteratively updating the hubness of the question based on the initial value of the hubness of the question and the latest value of the authority of each question to be recommended includes: For a target asked question, obtaining a transition probability from the target asked question to each question to be recommended; the target asked question is one of the asked questions in the set of asked questions; Obtaining the fourth product of the latest value of the authority of each question to be recommended and the transition probability of the target asked question to the corresponding question to be recommended; Summing the fourth products corresponding to the questions to be recommended to obtain a third summation calculation result; Obtaining a fifth product of the first weight coefficient and the third summation calculation result, and obtaining a sixth product of the second weight coefficient and the initial value of the hub degree of the target asked question; the sum of the first weight coefficient and the second weight coefficient is 1; The fifth product and the sixth product are summed to obtain a fourth sum calculation result, and the fourth sum calculation result is used as the hub degree of the target asked question after the iterative update.

4. The method according to claim 2 or 3, characterized in that The transition probability between two questions is obtained as follows: Traversing the historical text material set, counting the number of adjacent questions in each historical text material that are asked in sequence, and adding the count value to the corresponding position of the two adjacent questions in the first transfer matrix until the traversal is completed, thereby obtaining a second transfer matrix; the first dimension and the second dimension of the first transfer matrix both represent questions, and the elements at each position are set to 0. The first transfer matrix is ​​an initialization structure of the question transfer matrix; Normalizing the second transfer matrix according to the first dimension to obtain a problem transfer matrix; If it is necessary to obtain the transition probability from the first question to the second question, extract the value of the element corresponding to the first question in the first dimension and the second question in the second dimension in the question transfer matrix as the transition probability from the first question to the second question.

5. The method according to claim 4, characterized in that If the latest question in the set of asked questions is not a question extracted from the historical text material library, the transfer probability from the latest question to the target question to be recommended is calculated as follows: Calculating the similarity between the question asked the latest and other questions asked in the set of questions; Extracting the transition probabilities of other asked questions to the target question to be recommended from the question transition matrix; Multiply the similarities between the latest asked question and other asked questions in the set of asked questions, the transition probabilities of each of the other asked questions to the target question to be recommended, and the attenuation coefficients according to the corresponding values ​​of the other asked questions, and sum them up to obtain a fifth summation result as the preliminary transition probability from the latest asked question to the target question to be recommended; The fifth summation calculation result is divided by the sum of the preliminary transition probabilities for normalization, and the calculated quotient is used as the normalized transition probability from the question asked at the latest question time to the target question to be recommended; the sum of the preliminary transition probabilities is the sum of the preliminary transition probabilities from the question asked at the latest question time to each question to be recommended.

6. The method according to claim 1, characterized in that Initializing the hub degree of each question based on the time of each question asked and the similarity between multiple known questions includes: The multiple known questions are used as nodes in a question association graph, and according to the order of the questioning time of each question, a time decay weight of each question is determined as the initial point weight of the corresponding node in the question association graph; Using the similarities between any two of the plurality of known questions as edge weights between two corresponding nodes in the question association graph; The way to initialize the hub degree of the target asked question is: For a target node in the question association graph, determine a set of nodes that directly point to the target node in the question association graph, further obtain a question weight for each node in the node set, determine an edge weight from each node in the node set to the target node, and obtain an out-degree for each node in the node set; the question corresponding to the target node is the target asked question; and the target asked question is one of the asked questions in the set of asked questions; Multiply the problem weight of the same node in the node set by the edge weight from the node to the target node, and divide the result by the out-degree of the node to obtain the calculation result corresponding to the node; Adding the calculation results corresponding to each node in the node set to obtain the problem weight of the target node; The question weight of the target node is multiplied by the initial point weight of the target node to obtain the initial value of the hub degree of the target asked question.

7. The method according to claim 1, characterized in that Initializing the authority of each question to be recommended based on the recommendation score of each question to be recommended includes: Determine the minimum score and the maximum score among the recommended scores of each question to be recommended; Obtaining a first difference between the maximum score and the minimum score; For each question to be recommended, obtaining a second difference between the recommendation score of the question to be recommended and the minimum score; The ratio of the second difference to the first difference is used as the initial value of the authority of the question to be recommended.

8. The method according to claim 1 or 6, characterized in that The similarity between the first question and the second question is calculated as follows: Obtain a word segmentation set for the first question and a word segmentation set for the second question; Calculate the number of words in the intersection of the word set of the first question and the word set of the second question, and calculate the number of words in the union of the word set of the first question and the word set of the second question; The ratio of the number of the intersection words to the number of the union words is used as the similarity between the first question and the second question.

9. The method according to claim 1, characterized in that The set of questions to be recommended is in the form of a list of questions to be recommended. In the list of questions to be recommended, each question to be recommended is arranged in descending order according to the recommendation score. The list of questions to be recommended is constructed as follows: Using the keywords and word frequencies of the target text material, a plurality of feature vectors are retrieved from a pre-constructed co-occurrence matrix; the feature vectors correspond one-to-one to the keywords of the target text material; the co-occurrence matrix is ​​a two-dimensional matrix, wherein the first dimension represents questions and the second dimension represents keywords; wherein the questions represented by the first dimension of the co-occurrence matrix are question sentences extracted from a collection of historical text materials, and the keywords represented by the second dimension of the co-occurrence matrix are text keywords extracted from the collection of historical text materials; Performing feature enhancement processing on feature vectors corresponding to corresponding keywords using the word frequencies to obtain multiple enhanced feature vectors; Performing a sum operation on the multiple enhanced feature vectors to obtain a result vector; A list of to-be-recommended questions for the target text material is generated according to the recommendation scores in the result vector.

10. The method according to claim 1, characterized in that After the iteration is completed, a question recommendation result is generated based on the authority of each question to be recommended, including: After the iteration is completed, the top X recommended questions with higher authority and their corresponding authority are included in the question recommendation result, and the question recommendation result is output; or after the iteration is completed, the top Y% recommended questions with higher authority and their corresponding authority are included in the question recommendation result, and the question recommendation result is output; X is a positive integer, and X is less than the total number of questions to be recommended in the set of questions to be recommended; Y is a positive number less than 100.

11. The method according to claim 1, wherein After generating the question recommendation results according to the authority of each question to be recommended, the method further includes: Recording the next question selected in the question recommendation result; The next question is deleted from the set of questions to be recommended, and the next question is included in the set of questions that have been asked.

12. A question dynamic recommendation device based on HITS algorithm, characterized in that: include: A question acquisition module is configured to acquire a set of asked questions and a set of questions to be recommended; wherein each asked question in the set of asked questions is abstracted as a hub page in the HITS algorithm, and each question to be recommended in the set of questions to be recommended is abstracted as an authoritative page in the HITS algorithm; A first initialization module is configured to initialize the hubness of each question asked based on the time of each question asked and the similarity between multiple known questions; the multiple known questions include at least each question asked and each question to be recommended; The second initialization module is used to initialize the authority of each question to be recommended according to the recommendation score of each question to be recommended; A first iteration module is configured to iteratively update the authority of each question to be recommended based on an initial value of the authority of the question to be recommended and the latest value of the hub degree of each question that has been asked; A second iteration module is configured to iteratively update the hub degree of each question that has been asked based on the initial value of the hub degree of the question and the latest value of the authority of each question to be recommended; The recommendation module is used to generate question recommendation results based on the authority of each question to be recommended after the iteration is completed.

13. A question dynamic recommendation device based on HITS algorithm, characterized in that: include: memory and processor; The memory is used to store computer programs; The processor is configured to run the computer program, and when the computer program is run, the steps of the question dynamic recommendation method based on the HITS algorithm are executed as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method and device for recommending question and answer page related questions

    CN104462554A

  • A method for recommending points of interest based on the importance of the points of interest and the authoritativeness of a user

    CN109190053A