An Academic Topic Discovery Method Based on Reinforcement Learning

Through a method based on reinforcement learning, academia topic query representation and reasoning path is constructed, and a time-sequence convolution network and candidate action neighborhood encoder are used to solve the problem of academic topic discovery caused by the explosive growth of scientific and technological literature, and accurate prediction of popular academic topics that may appear in the future.

CN116860918BActive Publication Date: 2025-06-27CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310858312.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-13
Publication Date
2025-06-27
Estimated Expiration
2043-07-13

AI Technical Summary

Technical Problem

In the era of big data, the number of scientific and technological literature has exploded, making it difficult for scientific researchers to quickly and accurately discover potential academic topics and research trends, and existing technologies are difficult to predict popular academic topics that may appear in the future.

Method used

Using reinforcement learning-based method, by constructing academic topic query representation and inference paths, using a time-series convolution network and candidate action neighborhood encoder, we calculate the probability that the action is selected, select the execution action and update the inference path until the maximum inference step is reached, and the final path is evaluated based on the reward function, and the parameters are optimized to achieve academic topic discovery.

Benefits of technology

By integrating entity, relationship, and time characteristics, the one-sided nature of the cognitive actions of the agent is reduced, the accuracy of academic topic discovery is improved, and popular academic topics that may appear in the future can be predicted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116860918B_ABST
    Figure CN116860918B_ABST
Patent Text Reader

Abstract

The present invention discloses an academic topic discovery method based on reinforcement learning. The steps include: Step S1: Construct an academic topic query representation including entities, relationships, and time; Step S2: Model the inference path and action representation; Step S3: Calculate the correlation between the query representation and the actions and filter the actions. Further, score the actions based on the state of the inference path, and calculate the probability of an action being selected by combining the correlation and the score; Step S4: Select and execute an action based on the probability, update the inference path, stop when reaching the answer entity or the maximum number of steps, and optimize the parameters by maximizing the expected cumulative reward to obtain an academic topic discovery model and realize academic topic discovery. The advantages are that the present invention constructs a query representation that integrates entities, relationships, and time, regards entities, relationships, and time as an associated whole, and makes the query representation have more complete semantics; in addition, the present invention constructs an action representation by combining action neighborhood facts, reducing the one-sidedness of the understanding of actions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and particularly relates to a method for discovering academic topics based on reinforcement learning. Background Art

[0002] For scientific researchers, closely following the current latest academic hotspots and cutting-edge trends from scientific and technological literatures, continuously discovering new problems and proposing new methods is the main way to maintain academic innovation. However, in today's big data era, with the rapid development of various disciplinary technologies, the number of scientific and technological literatures has grown explosively. Web of Science includes more than 20,000 journals and conferences in the fields of computer science, social sciences, engineering technology, etc. Today's various literature databases provide rich materials for the majority of scientific researchers to select research directions and understand cutting-edge trends. However, while large-scale and dynamically updated scientific and technological literatures meet the information needs of scientific researchers, they also bring serious pressure of information overload. How to obtain the evolution trend of research topics in specific fields from a wide variety of scientific research fields and a large number of scientific and technological literatures, grasp the context of scientific research development, and quickly and accurately discover potential topics has become a common problem for scientific researchers.

[0003] Traditional topic discovery methods are to use natural language processing models to mine the text of existing papers, calculate the popularity of each topic, and output the final popular topics. However, such methods summarize the current popular topics based on the text and cannot accurately predict the popular topics in the future.

[0004] In summary, there is an urgent need for a method for discovering academic topics based on reinforcement learning to predict the possible popular academic topics in the future to solve the problems existing in the prior art. Summary of the Invention

[0005] The object of the present invention is to provide a method for discovering academic topics based on reinforcement learning, and the specific technical solution is as follows:

[0006] A method for discovering academic topics based on reinforcement learning, the steps are as follows:

[0007] Step S1: Construct an academic topic query representation, represent the academic topic query based on a temporal convolutional network model to obtain an academic topic query representation, where the academic topic query includes a head entity, a relationship, and a time;

[0008] Step S2: Construct an inference path and an action set, model the inference path and the action representation, where the inference path is a sequence of actions taken by the agent starting from the academic topic query representation; the action set is an out-edge set sampled according to the current entity on the knowledge graph in the order from near to far in time, and the action is represented by a candidate action neighborhood encoder;

[0009] Step S3: Calculate the probability of each action being selected. Specifically, calculate the correlation between the academic topic query representation and the action representation, filter the actions in the action set based on the correlation to obtain a refined action set, score the refined action set in combination with the inference path, and balance the correlation and the current score through weights to obtain the final action score, that is, the probability of each action being selected;

[0010] Step S4: Select an execution action based on the probability of each action being selected, update the inference path until the maximum inference step is reached, evaluate the final path based on the reward function, and optimize the parameters according to the maximized cumulative reward expectation to obtain an academic topic discovery model, and perform prediction based on the academic topic discovery model to achieve academic topic discovery.

[0011] Specifically, in step S1, the head entity is the journal name, and the relationship is the academic topic.

[0012] Specifically, in step S1, representing the academic topic query means learning a deep feature for each academic topic query. The specific process is as follows:

[0013] The first step is to define the time vector t i :

[0014] t i = cos(w t t i + b t );

[0015] Among them, w t , b t ∈ d t , both of which are learnable parameters; t i represents time; cos represents the cosine activation function.

[0016] The second step is to construct the entity embedding and the relationship embedding The construction formulas are as follows:

[0017]

[0018] Among them, h represents the entity or relationship. The first equation captures the temporal features of the first γd e elements before the entity embedding, and the second equation captures the static features;

[0019] The third step is to perform a one-dimensional convolution operation. The expression is as follows:

[0020]

[0021] M C = [m1, m2,..., m C;

[0022] Among them, represents the nth element of the output vector m of the convolutional channel c; c = 1, 2,..., C, representing the output channels; w c ∈ c is the learnable parameter matrix of the output channel c; C represents the number of output channels; K represents the convolutional kernel width; M C represents the new feature constructed by C convolutional channels; C The final output feature of the temporal convolutional network model is the academic topic query representation, which is defined as follows:

[0023] q = TCTE(e q , r q , t q ) = f(M C W1 + b1);

[0025] Among them, W1 ∈ Cd×d , b1 ∈ b , both of which are learnable parameters; f(·) represents the non-linear activation function; TCTE(·) represents the temporal convolutional network model; q is the vector representation of the academic topic query (e q , r q , t q ).

[0026] Specifically, in step S2, the inference path is constructed through the LSTM model, and the expression is as follows:

[0027] p l = LSTM(p l-1 , a l );

[0028] Among them, a l is the action taken in the l-th step decision; p l-1 represents the inference path before executing the action a l ; p l represents the inference path after executing the action a l .

[0029] Specifically, in step S2, the action set includes a self-loop action. When the self-loop action is executed, the state value of the LSTM is not updated.

[0030] Specifically, in step S2, the expression of the candidate action neighborhood encoder is as follows:

[0031]

[0032]

[0033] Among them, a iIt is the representation of action a i which includes the characteristics of the action itself and neighborhood characteristics; Representing action a i the entity e in i at time t i the set of actions, where W and b represent learnable parameters.

[0033] Specifically, in step S3, the calculation expression of the correlation is as follows:

[0034] S q = A l q;

[0035] S q = S q + inf * f l ;

[0036] Among them, represents the refined correlation; A l ∈ N×d , representing the action space, and the action space represents the embedding matrix composed of all action representations; q represents the academic topic query representation; S q ∈ N×1 , representing the matrix composed of the correlation coefficients between actions and queries; inf is a large negative number; f l ∈ N×1 represents the screening array, and the specific construction method of the screening array is as follows:

[0037] To achieve the screening of actions, record the score topk and the array f l ∈ N×1 , the value of the j-th position in the array is 1, indicating that the score at the position in the scoring matrix Q scores is less than topk, otherwise it is 0; topk represents the k-th largest number in S q ; After screening, keep the k that are more relevant to the academic topic query.

[0038] Specifically, in step S3, the expression of the final action score S f is as follows:

[0039] S p = A l p l-1 ;

[0040] S f = βS p +(1 - β)S q ;

[0041] Among them, S p ∈ N×1 , representing the matrix composed of the correlation coefficients between actions and the inference state; β is the weight coefficient.

[0042] Specifically, in step S4, the reward function is defined as follows:

[0043] R(p L ) = I{e L ==e te};

[0044] Among them, R(P L ) represents the reward function; P L represents the final path; e L represents the entity reached by the maximum number of inference steps; e te represents the answer entity of the query.

[0045] Specifically, in step S4, the process of maximizing the expected cumulative reward to optimize the parameters is as follows:

[0046] Based on the final path P L and the expected reward of all academic topic queries in the maximized training set G train , training is carried out, and the expression is as follows:

[0047]

[0048] Among them, J(θ) represents the expected reward obtained in the training set G train ; represents the expected reward of the academic topic query (e q , r q , t q ); e q represents the query entity; r q represents the query relationship; t q represents the query time; R represents the reward function;

[0049] Based on the REINFORCE algorithm, the parameter θ is iteratively updated, and the approximate gradient calculation formula is as follows:

[0050]

[0051] Among them, represents the approximate gradient of the expected reward, and π θ (a l |s l ) represents the probability that the action performed in the l-th step occurs.

[0052] Applying the technical solution of the present invention has the following beneficial effects:

[0053] (1) The temporal convolutional network model in the present invention integrates entity, relationship, and time features at the overall level to obtain the feature vectors of queries and actions. Regarding entities, relationships, and time as an internally interconnected whole enables the obtained feature vectors to have better semantic features.

[0054] (2) The candidate action neighborhood encoder in the present invention calculates the features of the action space connected to each action, reduces the one-sidedness of the agent's cognitive actions, and improves the reasoning accuracy.

[0055] (3) The present invention screens the action representations in the action set by calculating the correlation between the academic topic query representation and the action representation, removes irrelevant actions, and further combines the reasoning state for scoring, improving the accuracy of academic topic discovery.

[0056] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The following will refer to the drawings for a further detailed description of the present invention. Brief Description of the Drawings

[0057] The drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0058] Figure 1 is the step flowchart of the academic topic discovery method in the preferred embodiment of the present invention;

[0059] Figure 2 is the data graph of the experimental case in the preferred embodiment of the present invention;

[0060] Figure 3 is the detailed data graph of the ablation experiment in the preferred embodiment of the present invention. Detailed Description of the Embodiment

[0061] The following will describe the embodiments of the present invention in detail with reference to the drawings, but the present invention can be implemented in various different ways defined and covered by the claims.

[0062] The method proposed by the present invention can be regarded as a Markov decision process (MDP). Specifically, through the interaction between the agent and the environment, the reasoning path to reach the answer entity is found. Specifically, the agent starts from the entity e q at time t q and transfers to a new state node e l until the maximum number of reasoning steps L or the answer node is reached.

[0063] Furthermore, the present invention designs a two-stage decision-making policy network π θ (a l+1 |sl ) = P(a l+1 |s l ; θ), which guides the agent to make decisions according to the current state s l and the inference path p l = (e q , r q , t q , a1,..., a l ), where a l ∈ A l represents the actions taken at the inference steps l = 1,..., L, and θ are the parameters of the model. It should be noted that p0 = (e q , r q , t q ).

[0064] Example:

[0065] See Figure 1 , this example discloses a method for discovering academic topics based on reinforcement learning, and the steps are as follows:

[0066] Step S1: Construct an academic topic query, and represent the academic topic query based on a Temporal ConvTransE model to obtain an academic topic query representation, where the academic topic query includes a head entity, a relationship, and a time;

[0067] Step S2: Construct an inference path and an action set, and model the inference path and action representation. Among them, the inference path is a sequence of actions taken by the agent starting from the academic topic query representation, and the inference path is modeled by an LSTM encoder; the action set is an out-edge set sampled from the knowledge graph according to the current entity in chronological order from near to far, and the action set includes the actions that appear in the inference path, and the actions are represented by a candidate action neighborhood encoder;

[0068] Step S3: Calculate the probability of each action being selected. Specifically, calculate the correlation between the academic topic query representation and the action representation, filter the actions in the action set based on the correlation to obtain a refined action set, score the refined action set in combination with the inference path, and balance the correlation and the current score through weights to obtain the final action score, that is, the probability of each action being selected;

[0069] Step S4: Select and execute actions based on the probability of each action being selected, update the inference path until the maximum inference step is reached, evaluate the final path based on the reward function, and optimize the parameters according to the maximization of the cumulative reward expectation to obtain an academic topic discovery model, and perform predictions based on the academic topic discovery model to achieve academic topic discovery.

[0070] It should be noted that the purpose of the Temporal ConvTransE model is to imitate the process of human beings' understanding of problems from a global perspective. Specifically, the problem is considered as a whole, rather than being decomposed into individual unrelated parts, so as to better understand the essence and structure of the problem.

[0071] Further, in step S1, the head entity in this embodiment is the journal name, and the relationship is the academic theme.

[0072] Further, in step S1, this embodiment uses the Temporal ConvTransE model to represent the academic theme query, that is, to learn a deep feature for each academic theme query. The specific process is as follows:

[0073] The first step is to define the time vector t i :

[0074] t i = cos(w t t i + b t );

[0075] Among them, w t , b t ∈ d t , both of which are learnable parameters; t i represents time. In addition, entities and relationships may have time-varying features, while some features may be static. Therefore, this embodiment further constructs entity embeddings and relationship embeddings

[0076] The second step is to construct entity embeddings and relationship embeddings The construction formula is as follows:

[0077]

[0078] Among them, h represents an entity or a relationship. The first equation captures the temporal features of the first γd e elements before the entity embedding, and the second equation captures the static features; in order to reduce the semantic bias existing in the query (e i , r i , t i ) embedding, this embodiment further uses one-dimensional convolution to fuse the features of the entity relationship time t i ;

[0079] The third step is to perform one-dimensional convolution operation, and the expression is as follows:

[0080]

[0081] M C = [m1, m2,..., m C ;

[0082] Among them, represents the nth element of the output vector m c of the convolutional channel c; c = 1, 2,..., C represents the output channels; w c ∈ K×3 is the learnable parameter matrix of the output channel c; C represents the number of output channels; K represents the convolutional kernel width; M C represents the new feature constructed by C convolutional channels;

[0083] Finally, the final output feature of the Temporal ConvTransE model is the academic topic query representation, which is defined as follows:

[0084] q = TCTE(e q , r q , t q ) = f(M c W1 + b1);

[0085] Among them, W1 ∈ Cd×d and b1 ∈ b , both of which are learnable parameters; f(·) represents the non-linear activation function; TCTE(·) represents the temporal convolutional network model; q is the vector representation of the academic topic query (e q , r q , t q ).

[0086] Specifically, in step S2, the inference path p l is constructed through the LSTM model, and the expression is as follows:

[0087] p l = LSTM(p l-1 , a l );

[0088] Among them, a l is the action taken in the l-th step decision; p l-1 represents the inference path before executing the action a l ; p l represents the inference path after executing the action a l . It should be noted that p0 = LSTM(0, [r q , e q , t q);In addition, the action set in this embodiment includes self-loop actions. If a self-loop action is taken in the previous step, the state value of the LSTM is not updated.

[0089] Specifically, in step S2, the expression of the candidate action neighborhood encoder is as follows:

[0090]

[0091] where a i represents the neighborhood feature of action a i , that is, the action representation; represents the set of actions of entity e i in action a i at time t i ; f(·) is a non-linear activation function. The activation function preferred in this embodiment is the RELU activation function; b represents learnable parameters. The candidate action neighborhood encoder in this embodiment aggregates the information of neighborhood facts and reduces the one-sidedness in the process of constructing actions.

[0092] Specifically, in step S3, the calculation expression of the relevance is as follows:

[0093] S q = A l q;

[0094] S q = S q + inf*f l ;

[0095] where represents the refined relevance; A l ∈ N×d , is the action space, and the action space represents the embedding matrix composed of all actions; q represents the academic topic query representation; S q ∈ N×1 is the matrix composed of the correlation coefficients of actions and queries; inf is a large negative number; f l ∈ N×1 represents the screening array, and the specific construction method is as follows:

[0096] To achieve the screening of actions, record the score topk and the array f l ∈ N×1 . The value of the j-th position in the array is 1, indicating that the score at the position in the scoring matrix Q scores is less than topk, otherwise it is 0; topk represents the k-th largest number in S q ; After screening, keep the k actions that are more relevant to the academic topic query.

[0097] Specifically, in step S3, the final action score S fThe expression is as follows:

[0098] S p = A l p l-1 ;

[0099] S f = βS p + (1 - β)S q ;

[0100] Where S p ∈ N×1 , represents the matrix composed of the correlation coefficients of actions and reasoning states; β is the weight coefficient.

[0101] After scoring the action set A l , the π θ (a l+1 |s l ) can be obtained through the softmax activation function. After a total of L rounds of iteration, the inference path P = ((e q , r q , t q ), (e1, r1, t1),..., (e L , r L , t L )) can be obtained. e L is the finally discovered topic, and at the same time, the inference path P makes a certain explanation for the inference result e L , improving the user's trust in the model.

[0102] Specifically, in step S4, the definition of the reward function is as follows:

[0103] R(p L ) = I{e L == e te};

[0104] Where R(P L ) represents the reward function; P L represents the final path; e L represents the entity reached by the maximum number of inference steps; e te represents the answer entity of the query.

[0105] Specifically, in step S4, the process of maximizing the expected cumulative reward to optimize the parameters is as follows:

[0106] Based on the final path P L and the expected rewards represented by all academic topic queries in the maximized training set G train , the training is carried out. The expression is as follows:

[0107]

[0108] Among them, J(θ) represents the expected reward obtained from the training set G train ; represents the expected reward for the academic topic query (e q , r q , t q ); e q represents the query entity; r q represents the query relationship; t q represents the query time; R represents the reward function;

[0109] This embodiment iteratively updates the parameter θ based on the REINFORCE algorithm, and the approximate gradient calculation formula is as follows:

[0110]

[0111] Among them, represents the approximate gradient of the expected reward, and π θ (a l |s l ) represents the probability that the action executed at the l-th step occurs.

[0112] Simulation experiment:

[0113] I. Experimental setup

[0114] 1. Dataset

[0115] In the present invention, the academic domain knowledge graph datasets Aminer Topic Top 5000 Publications and Authors and DBLP - Paper - Topic are used, which respectively contain data on topics, publications, authors, and the topics described in each paper. The present invention sequentially selects 80%, 10%, and 10% of them as the training set, validation set, and test set in chronological order.

[0116] 2. Evaluation metrics

[0117] Two widely used metrics, Hits@K and Mean Reciprocal Rank (MRR), are used to measure the performance of the model in link prediction at future timestamps. Hits@K means the proportion of quadruples with a score rank less than or equal to K. MRR represents the average of the reciprocals of these ranks.

[0118] 3. Model parameters

[0119] The models in this experiment are all implemented using PyTorch and all experiments are executed on a 32G Tesla V100. For all datasets, the dimensions of the relationship, entity, and time embedding vectors are set to 100, 80, and 20 respectively. In this experiment, the ADAM optimizer is used to optimize the network parameters, and the learning rate is set to 0.0001. The batch size during training is set to 1024. The size N of the candidate action set is set to 50, and the inference path length is 2.

[0120] 4. Case

[0121] Users can trace back according to the inference path provided by the invention to find out why such an inference is made, which provides a factual basis for the appearance of the inference result. The query of this method in the academic field is explored and the results are shown in Figure 2 . Taking the first query in the figure (Computer Science and Technology, subject,?, 2022) as an example, in this experiment, considering that "robot control" is a relatively popular research field in the discipline of Computer Science and Technology in the relatively recent period, the agent transfers on the knowledge graph, from the discipline entity "Computer Science and Technology" to the domain entity "robot control", and then further considering that the subject entity "reinforcement learning" is a research method often used in the domain "robot control", so it finally selects the subject "reinforcement learning" as the answer output. Taking the domain as the head entity for subject discovery, taking (Computer Vision, subject,?, 2022) as an example, the agent first starts from the research domain entity "Computer Vision", goes through autonomous driving until its difficult subject "image segmentation", and finally takes the subject "image segmentation" as the answer entity.

[0122] 5. Ablation Experiment

[0123] To explore the contribution of each step of this method to the prediction performance, this experiment also conducts ablation experiments. As Figure 3As shown, the results of different variants of the present method are reported. To verify the contribution of the Temporal ConvTransE model to the prediction results, its implementation was removed in this experiment and replaced with a fully connected layer. It was replaced with a single fully connected layer. The results are shown as "w / o TCTE" in the figure. Obviously, the performance of w / o TCTE is not as good as that of the complete model. This is because the TemporalConvTransE model can model queries and actions more comprehensively and accurately, resulting in a more accurate and comprehensive understanding of the queries and actions related to the subject entity, discipline entity, and domain entity, proving the effectiveness of the proposed Temporal ConvTransE model. w / oAN in the figure shows the performance after removing the candidate action neighborhood encoder. It is not difficult to see from the figure that the performance of w / oAN is poor, which is because the candidate action neighborhood encoder models the neighborhood facts, helping to reduce the one-sidedness of the understanding of actions. Specifically, in the problem of academic topic discovery, there are cross, complementary, dependent, and mutually influential relationships between adjacent fields, adjacent topics, and adjacent disciplines. The present invention can more accurately understand the current actions taken through its neighborhood, improving the model performance. To prove the contribution of the two-stage action evaluation function to the final performance of the present method, this experiment removed stage one and stage two respectively. It is represented as w / o stage one and w / o stage two in Figure 3 respectively. Due to the lack of guidance from the query, the performance in w / o stage one decreased significantly. Due to the lack of scoring based on the current state, the performance of w / o stage two did not achieve the optimum either. Specifically, reasoning is a multi-step process. Therefore, not only should the current problem be clarified based on the query, but also the information inferred currently, that is, the current reasoning state, needs to be analyzed and evaluated. Therefore, it is inappropriate to rely solely on the query or the current state for academic topic discovery. Actions should be taken for reasoning by combining the query and the current state simultaneously. Therefore, the present invention adopts a two-stage action evaluation function to determine the action selection.

[0124] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An academic theme discovery method based on reinforcement learning, characterized in that The steps are as follows: Step S1: Construct an academic topic query representation. Based on a temporal convolutional network model, represent the academic topic query to obtain an academic topic query representation. The academic topic query includes a head entity, a relationship, and a time; Step S2: Construct an inference path and an action set, and model the inference path and action representation. Among them, the inference path is a sequence of actions taken by the agent starting from the academic topic query representation; the action set is an out-edge set sampled from the knowledge graph according to the current entity in chronological order from near to far, and the action is represented by a candidate action neighborhood encoder; Step S3: Calculate the probability of each action being selected. Specifically, calculate the correlation between the academic topic query representation and the action representation, filter the actions in the action set based on the correlation to obtain a refined action set, score the refined action set in combination with the inference path, and balance the correlation and the current score through weights to obtain the final action score, that is, the probability of each action being selected; Step S4: Select and execute actions based on the probability of each action being selected, update the inference path until the maximum inference step is reached, evaluate the final path based on the reward function, and optimize the parameters according to the maximized cumulative reward expectation to obtain an academic topic discovery model. Perform predictions based on the academic topic discovery model to achieve academic topic discovery; In step S1, the head entity is the journal name, and the relationship is the academic topic; In step S1, representing the academic topic query is to learn a deep feature for each academic topic query. The specific process is as follows: Step 1, define the time vector t i : t i = cos(w t t i + b t ); where, w t , both are learnable parameters; t i represents time; cos represents the cosine activation function; Step 2: Construct entity embeddings and relationship embeddings respectively and relationship embeddings The construction formula is as follows: Among them, h represents an entity or a relationship. The first equation captures the temporal features of the first γd e elements before entity embedding, and the second equation captures the static features; The third step is to perform a one-dimensional convolution operation. The expression is as follows: M C = [m1, m2,..., m C ; Among them, represents the n-th element of the output vector m of the convolutional channel c; c c = 1, 2, ..., C, representing the output channels; is the learnable parameter matrix of the output channel c; C represents the number of output channels; K represents the convolutional kernel width; M C represents the new feature constructed by C convolutional channels; The final output feature of the temporal convolutional network model is the academic topic query representation, which is defined as follows: q = TCTE(e q , r q , t q ) = f(M C W1 + b1); Among them, both are learnable parameters; f(·) represents a non-linear activation function; TCTE(·) represents a temporal convolutional network model; q is the vector representation of the academic topic query (e q , r q , t q ).

2. The academic topic discovery method according to claim 1, wherein In step S2, the inference path is constructed through an LSTM model. The expression is as follows: p l = LSTM(p l-1 , a l ); where a l is the action taken in the l-th step decision; p l-1 represents the reasoning path before executing action a l ; p l represents the reasoning path after executing action a l .

3. The academic topic discovery method according to claim 2, wherein In step S2, the action set includes a self-loop action. When the self-loop action is executed, the state value of the LSTM is not updated.

4. The academic topic discovery method according to claim 3, characterized in that In step S2, the expression of the candidate action neighborhood encoder is as follows: Among them, a i is the representation of action a i which includes the characteristics of the action itself and neighborhood characteristics; represents the set of actions of entity e i in action a i at time t i where W and b represent learnable parameters.

5. The academic topic discovery method according to claim 4, wherein In step S3, the calculation expression of the correlation is as follows: S q = A l q; Among them, represents the refined relevance; represents the action space, and the action space represents the embedding matrix composed of all action representations; q represents the academic topic query representation; represents the matrix composed of the correlation coefficients of actions and queries; inf is a large negative number; represents the screening array, and the specific construction method of the screening array is as follows: To achieve the screening of actions, record the score topk and the array The value of the corresponding position j in the array is 1, indicating that the score at the position in the rating matrix Q scores is less than topk, otherwise it is 0; topk represents the k-th largest number in S q After screening, retain the k ones that are more relevant to the academic theme query.

6. The academic topic discovery method according to claim 5, characterized in that In step S3, the final action score S f has the following expression: S p = A l p l-1 ; Among them, represents the matrix formed by the correlation coefficients of actions and inference states; β is the weight coefficient.

7. The academic topic discovery method according to claim 6, wherein In step S4, the definition of the reward function is as follows: Among them, R(P L ) represents the reward function; P L represents the final path; e L represents the entity reached at the maximum number of inference steps; e te represents the answer entity of the query.

8. The academic topic discovery method according to claim 7, wherein In step S4, the process of optimizing the parameters according to the maximized cumulative reward expectation is as follows: Based on the final path P L and maximizing the expected reward for all academic topic queries in the training set G train is trained with the following expression: Among them, J(θ) represents the expected reward obtained from the training set G train ; represents the expected reward of the academic topic query (e q , r q , t q ); e q represents the query entity; r q represents the query relationship; t q represents the query time; R represents the reward function; Iteratively update the parameter θ based on the REINFORCE algorithm. The approximate gradient calculation formula is as follows: Among them, represents the approximate gradient of the expected reward, π θ (a l |s l ) represents the probability that the action performed at the l-th step occurs.

Citation Information

Patent Citations

  • Processing search queries for open education resources

    US20170011095A1

  • Systems and methods for human inspired simple question answering (HISQA)

    US20170109355A1