A journal matching and recommendation method and device based on big data

Through machine learning models and knowledge graphs based on feature vectors, combined with similarity and matching, the journal recommendation method is optimized, and the problems of inaccurate labels and large amounts of calculations in the existing technology are solved, and efficient and accurate journal matching is achieved.

CN119357468BActive Publication Date: 2025-07-22GUANGZHOU KEAO INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411401798.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-09
Publication Date
2025-07-22
Estimated Expiration
2044-10-09

AI Technical Summary

Technical Problem

The existing journal matching methods rely on the accuracy of label settings, resulting in inaccurate output, and the calculation is large and time-consuming when calculating similarity.

Method used

By obtaining the feature vectors of the manuscripts and journals to be detected, training the machine learning model, combining similarity and matching degree for journal recommendations, and using emotional feature vectors, word frequency feature vectors, and knowledge graphs to calculate the popularity scores, and optimizing the recommendation sorting.

Benefits of technology

It improves the accuracy and speed of journal matching, reduces the error of artificial label settings, and reduces the amount of data calculated in similarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357468B_ABST
    Figure CN119357468B_ABST
Patent Text Reader

Abstract

This application belongs to the field of computer technology and discloses a periodical matching and recommendation method and device based on big data. The method includes: obtaining a manuscript to be detected and extracting the feature information therein to form a target feature vector; obtaining historical periodical data and extracting the feature information of each periodical to form a corresponding periodical feature vector; training a machine learning model based on each periodical feature vector to obtain a periodical matching model; inputting the target feature vector into the periodical matching model to obtain multiple matching periodicals and corresponding matching degrees; calculating the similarity between the target feature vector and the periodical feature vectors of each matching periodical; and sorting each matching periodical based on the similarity and the matching degree to obtain a recommended periodical list. This application can reduce the influence of human subjective factors in model matching and the computational amount required for similarity calculation, and greatly improve the accuracy and speed of periodical matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method and device for periodical matching and recommendation based on big data. Background Art

[0002] In current various periodical statistical databases or paper management systems, it is a common function to intelligently match relevant periodical data according to the documents uploaded by users for researchers to consult and study. The current intelligent matching methods mainly include model matching based on various pre-set tags (such as keywords, etc.) on each periodical, or matching methods based on text similarity.

[0003] However, traversing tags through a large model to match relevant periodicals highly depends on the accuracy of tag settings. When the keywords set by periodical authors are inappropriate, it will lead to inaccurate periodical tags and affect the accuracy of model output; the method of calculating similarity requires processing a large amount of periodical data, resulting in long calculation time and high computing power consumption. Summary of the Invention

[0004] This application provides a method and device for periodical matching and recommendation based on big data, which can reduce the influence of human subjective factors in model matching and greatly improve the accuracy and speed of periodical matching.

[0005] In a first aspect, an embodiment of this application provides a method for periodical matching and recommendation based on big data, including:

[0006] Obtain the manuscript to be detected, and extract the feature information therein to form a target feature vector;

[0007] Obtain historical periodical data, and extract the feature information of each periodical to form a corresponding periodical feature vector;

[0008] Train a machine learning model based on each periodical feature vector to obtain a periodical matching model;

[0009] Input the target feature vector into the periodical matching model to obtain multiple matching periodicals and corresponding matching degrees;

[0010] Calculate the similarity between the target feature vector and the periodical feature vectors of each matching periodical;

[0011] Sort each matching periodical based on the similarity and the matching degree to obtain a recommended periodical list.

[0012] Furthermore, the method further includes:

[0013] Before extracting the feature information of the manuscript to be detected, remove the preset stop words in the manuscript to be detected;

[0014] Perform word segmentation, stemming, and lowercase conversion on the manuscript to be detected.

[0015] Further, obtaining the document to be detected and extracting the feature information therein to form a target feature vector includes:

[0016] Extracting the sentiment feature vector of the document to be detected based on a preset sentiment dictionary;

[0017] Extracting the word frequency feature vector of the document to be detected using the TF-IDF method;

[0018] Combining the sentiment feature vector and the word frequency feature vector to obtain the target feature vector.

[0019] Further, extracting the sentiment features of the document to be detected based on a preset sentiment dictionary includes:

[0020] Assigning sentiment polarity scores to each word in the document to be detected according to the preset sentiment dictionary;

[0021] Calculating the document sentiment score of the document to be detected based on each sentiment polarity score;

[0022] Determining the sentiment type and sentiment intensity of the document to be detected according to the document sentiment score;

[0023] Constructing the document sentiment score, sentiment type and sentiment intensity into a sentiment feature vector.

[0024] Further, the method further includes:

[0025] Before training the machine learning model, calculating the chi-square statistic of each journal feature vector;

[0026] Constructing a contingency table corresponding to each journal feature vector; wherein, the rows of the contingency table represent the values of different features in the journal feature vector, and the columns represent different feature categories;

[0027] Determining the chi-square threshold of the journal feature vector according to the preset significance level and the contingency table;

[0028] If the chi-square statistic is less than the chi-square threshold, then eliminating the corresponding journal feature vector.

[0029] Further, determining the chi-square threshold of the journal feature vector according to the preset significance level and the contingency table includes:

[0030] Let the number of rows and columns of the contingency table be each decreased by 1, and then multiply to obtain the degrees of freedom;

[0031] Let the preset significance level be multiplied by the degrees of freedom to obtain the chi-square threshold.

[0032] Further, calculating the similarity between the target feature vector and the journal feature vectors of each matching journal includes:

[0033] Calculate the Manhattan distance or Jaccard similarity coefficient between the target feature vector and the journal feature vector, and use the Manhattan distance or Jaccard similarity coefficient as the similarity.

[0034] Further, sort each matching journal based on the similarity and matching degree to obtain a recommended journal list, including:

[0035] Normalize each similarity and each matching degree respectively;

[0036] Let the normalized similarity and matching degree be added to obtain the sorting value of the matching journal;

[0037] Sort each matching journal according to the sorting value to obtain a recommended journal list.

[0038] Further, the method further includes:

[0039] Calculate the popularity score of each matching journal based on the knowledge graph method to obtain a popularity sorted list;

[0040] Adjust the sorting of each matching journal in the recommended journal list according to the popularity sorted list.

[0041] Further, the above-mentioned method of calculating the popularity score of each matching journal based on the knowledge graph to obtain a popularity sorted list includes:

[0042] Take each matching journal as a node, take the journal feature vector as the attribute of the corresponding node, and take the citation relationship between each matching journal as the edge connecting each node to obtain a journal knowledge graph;

[0043] Obtain the comment popularity score according to the comment text and the corresponding comment time of each node;

[0044] Obtain the collection popularity score according to the cumulative collection volume and collection growth rate of each node;

[0045] Perform weighted calculation on the comment popularity score and the collection popularity score to obtain the node score;

[0046] Use the graph algorithm to calculate the popularity score of each node according to the node scores and the edges between the nodes;

[0047] Sort the matching journals corresponding to each node according to the popularity score to obtain a popularity sorted list.

[0048] Further, obtaining the collection popularity score according to the cumulative collection volume and collection growth rate of each node includes:

[0049] Determine the first score in the collection interval where the cumulative collection volume is located, and the second score in the growth rate interval where the collection growth rate is located;

[0050] Add the first score and the second score to obtain the collection popularity score.

[0051] Furthermore, adjusting the sorting of each matching journal in the recommended journal list according to the popularity sorted list includes:

[0052] Take the average of the sorting of the matching journal in the popularity sorted list and the recommended journal list as the new sorting of the matching journal in the recommended journal list.

[0053] Furthermore, the method further includes:

[0054] Obtain the click times of each matching journal, and modify the corresponding matching degree based on the click times;

[0055] Train the journal matching model based on the modified matching degree and the corresponding journal feature vector.

[0056] In a second aspect, an embodiment of the present application provides a journal matching and recommending device based on big data, including:

[0057] A target manuscript module, configured to obtain a manuscript to be detected, and extract the feature information therein to form a target feature vector;

[0058] A journal module, configured to obtain historical journal data, and extract the feature information of each journal to form a corresponding journal feature vector;

[0059] A training module, configured to train a machine learning model based on each journal feature vector to obtain a journal matching model;

[0060] A matching module, configured to input the target feature vector into the journal matching model to obtain a plurality of matching journals and corresponding matching degrees;

[0061] A calculation module, configured to calculate the similarity between the target feature vector and the journal feature vectors of each matching journal;

[0062] A recommendation module, configured to sort each matching journal based on the similarity and the matching degree to obtain a recommended journal list.

[0063] Furthermore, the device further includes a preprocessing module, configured to remove preset stop words in the manuscript to be detected before extracting the feature information of the manuscript to be detected; and perform word segmentation, stemming, and lowercase conversion on the manuscript to be detected.

[0064] Furthermore, the above-mentioned target manuscript module includes:

[0065] An emotion unit, configured to extract the emotion feature vector of the manuscript to be detected based on a preset emotion dictionary;

[0066] A term frequency unit, configured to extract the term frequency feature vector of the manuscript to be detected by using the TF-IDF method;

[0067] A combination unit for combining the sentiment feature vector and the word frequency feature vector to obtain a target feature vector.

[0068] Furthermore, the above-mentioned sentiment unit is used for:

[0069] Assigning sentiment polarity scores to each word in the document to be detected according to a preset sentiment dictionary;

[0070] Calculating the document sentiment score of the document to be detected based on each sentiment polarity score;

[0071] Determining the sentiment type and sentiment intensity of the document to be detected according to the document sentiment score;

[0072] Constructing the document sentiment score, sentiment type and sentiment intensity into a sentiment feature vector.

[0073] Furthermore, the device further includes:

[0074] A statistical module for calculating the chi-square statistic of each journal feature vector before training the machine learning model;

[0075] A contingency table module for constructing a contingency table corresponding to each journal feature vector;

[0076] A threshold module for determining the chi-square threshold of the journal feature vector according to a preset significance level and the contingency table;

[0077] An elimination module for eliminating the corresponding journal feature vector if the chi-square statistic is less than the chi-square threshold.

[0078] Furthermore, the above-mentioned threshold module includes:

[0079] A degree of freedom unit for subtracting 1 from the number of rows and columns of the contingency table respectively and then multiplying to obtain the degree of freedom;

[0080] A threshold unit for multiplying the preset significance level and the degree of freedom to obtain the chi-square threshold.

[0081] Furthermore, the above-mentioned calculation module is used for calculating the Manhattan distance or Jaccard similarity coefficient between the target feature vector and the journal feature vector, and taking the Manhattan distance or Jaccard similarity coefficient as the similarity.

[0082] Furthermore, the above-mentioned recommendation module includes:

[0083] A normalization unit for normalizing each similarity and each matching degree respectively;

[0084] A sorting value unit for adding the normalized similarity and matching degree to obtain the sorting value of the matching journal;

[0085] A recommended list unit for sorting each matching journal according to a sorting value to obtain a recommended journal list.

[0086] Furthermore, the device further includes:

[0087] A popularity sorting module for calculating the popularity scores of each matching journal based on a knowledge graph method to obtain a popularity sorted list;

[0088] A sorting update module for adjusting the sorting of each matching journal in the recommended journal list according to the popularity sorted list.

[0089] Furthermore, the above-mentioned popularity sorting module includes:

[0090] A graph construction unit for taking each matching journal as a node, taking the journal feature vector as the attribute of the corresponding node, and taking the citation relationship between each matching journal as the edge connecting each node to obtain a journal knowledge graph;

[0091] A comment popularity unit for obtaining a comment popularity score according to the comment text of each node and the corresponding comment time;

[0092] A favorite popularity unit for obtaining a favorite popularity score according to the cumulative favorite volume and favorite growth rate of each node;

[0093] A node score unit for performing weighted calculation on the comment popularity score and the favorite popularity score to obtain a node score;

[0094] A node popularity unit for calculating the popularity scores of each node according to the node scores of each node and the edges between each node by using a graph algorithm;

[0095] A popularity sorting unit for sorting the matching journals corresponding to each node according to the popularity scores to obtain a popularity sorted list.

[0096] Furthermore, the above-mentioned favorite popularity unit is used for:

[0097] Determining a first score for the cumulative favorite volume in the favorite interval and a second score for the favorite growth rate in the growth interval;

[0098] Adding the first score and the second score to obtain a favorite popularity score.

[0099] Furthermore, the above-mentioned sorting update module is used to take the average of the sorting of the matching journal in the popularity sorted list and the recommended journal list as the new sorting of the matching journal in the recommended journal list.

[0100] Furthermore, the device further includes:

[0101] A click acquisition module for acquiring the click times of each matching journal and modifying the corresponding matching degree based on the click times.

[0102] A feedback training module for training a journal matching model based on the modified matching degree and the corresponding journal feature vector.

[0103] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it performs the steps of a big data-based journal matching and recommendation method according to any one of the above embodiments.

[0104] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of a big data-based journal matching and recommendation method according to any one of the above embodiments.

[0105] In summary, compared with the prior art, the beneficial effects brought by the technical solution provided by the embodiment of the present application at least include:

[0106] A big data-based journal matching and recommendation method provided by an embodiment of the present application first trains a journal matching model through journal feature vectors to obtain multiple matching journals and corresponding matching degrees, then calculates the similarity between the target feature vector of the manuscript to be detected and the journal feature vectors of each matching journal, and finally forms a recommended journal list by combining the similarity and the matching degree. The above method combines the similarity and the matching method of the large model based on feature vectors, which not only avoids the situation that the inaccurate labels set by humans affect the model output, greatly improves the accuracy of journal matching, but also reduces the amount of data required for similarity calculation, that is, the manuscript to be detected only needs to calculate the similarity with the matching journals output by the model, improving the matching speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] Figure 1 It is a flowchart of a big data-based journal matching and recommendation method provided by an exemplary embodiment of the present application.

[0108] Figure 2 It is a flowchart of the target feature vector construction step provided by an exemplary embodiment of the present application.

[0109] Figure 3 It is a flowchart of the emotional feature vector step provided by an exemplary embodiment of the present application.

[0110] Figure 4 It is a flowchart of the chi-square test step provided by an exemplary embodiment of the present application.

[0111] Figure 5 It is a flowchart of the journal ranking step provided by an exemplary embodiment of the present application.

[0112] Figure 6 Flowchart of the popularity ranking step provided for an exemplary embodiment of the present application.

[0113] Figure 7 Flowchart of the favorite popularity calculation step provided for an exemplary embodiment of the present application.

[0114] Figure 8 Flowchart of the model iteration step provided for an exemplary embodiment of the present application.

[0115] Figure 9 Structure diagram of a periodical matching and recommendation device based on big data provided for an exemplary embodiment of the present application.

[0116] Figure 10 Structure diagram of the target manuscript module provided for an exemplary embodiment of the present application.

[0117] Figure 11 Structure diagram of the recommendation module provided for an exemplary embodiment of the present application. Detailed implementation manners

[0118] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0119] All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0120] Term explanation:

[0121] 1) Similarity: It is an index that measures the closeness of two vectors in the feature space. Similarity is usually used to represent the similarity or difference between vectors, and it has wide applications in machine learning and data analysis, such as clustering analysis, recommendation systems, image recognition, etc. Similarity measurement methods include:

[0122] Cosine Similarity, which measures the similarity of two vectors by calculating the cosine value of the included angle between the two vectors. Cosine Similarity focuses on the direction of the vectors rather than their magnitudes.

[0123] Euclidean Distance, which measures the straight-line distance between two vectors. Contrary to Cosine Similarity, Euclidean Distance focuses on the magnitudes and directions of the vectors. The smaller the Euclidean Distance, the more similar the two feature vectors are.

[0124] The Manhattan Distance, also known as the L1 distance, is the sum of the absolute values of the differences in each dimension.

[0125] The Jaccard Similarity is used to measure the similarity between two sets and is calculated by the size of their intersection and union.

[0126] The Pearson Correlation Coefficient measures the linear correlation between two vectors.

[0127] The Spearman's Rank Correlation Coefficient is a non-parametric method based on ranks and is used to measure the monotonic association between two vectors.

[0128] The Mahalanobis Distance takes into account the spatial distribution of the data and uses the covariance matrix to calculate the distance.

[0129] The Levenshtein Distance measures the difference between two sequences by calculating the minimum number of single-character edits (insertions, deletions, or substitutions) required to transform one sequence into the other.

[0130] 2) Sentiment polarity score: It refers to the quantitative representation of the positive or negative tendency of a word in sentiment analysis. This score usually comes from a sentiment dictionary or is predicted by a large model and is used to judge the sentiment tendency expressed by each word in the text. The sentiment dictionary contains a large number of words and their corresponding sentiment polarity scores.

[0131] For example, AFINN, VADER (especially used for social media texts in the NLP field), and SentiWordNet are all commonly used sentiment dictionaries. Sentiment polarity is usually divided into positive, negative, and neutral. The polarity score usually varies between -1 (most negative) and 1 (most positive), and 0 usually represents neutral.

[0132] For a text, the sentiment tendency of the whole text can be calculated by summing or averaging the sentiment polarity scores of each word in the text. The sentiment polarity of a word may depend on the context. For example, "bank" usually has a positive meaning, but in some cases (such as "the bank goes bankrupt") it may have a negative meaning.

[0133] 3) Knowledge Graph: It is a structured semantic knowledge base that represents entities and the relationships between them in the form of a graph. The main features of the Knowledge Graph include:

[0134] Entities: Entities are the basic concepts in the Knowledge Graph and can be people, places, etc.

[0135] Relationships: Relationships define the semantic connections between entities, such as "belong to", "be located in", etc.

[0136] Attributes: Entities can have attributes that provide additional descriptive information about the entities, such as the date of birth of a person, the language of a country, etc.

[0137] Graph Structure: The Knowledge Graph adopts a graph form, where nodes represent entities and edges represent relationships, forming a complex network structure.

[0138] Multi-source Integration: The data of the Knowledge Graph can come from multiple different information sources, including structured data, semi-structured data, and unstructured data.

[0139] Dynamic Updating: The Knowledge Graph can be continuously updated and improved according to newly acquired data.

[0140] Scalability: The Knowledge Graph can be scaled to include millions or even billions of entities and relationships.

[0141] 4) Graph Algorithms: Graph algorithms refer to algorithms that run on graph-structured data and are used to perform various tasks, such as searching, analyzing, and mining information in the Knowledge Graph. The main graph algorithms for calculating node importance include the following:

[0142] Degree Centrality: Degree Centrality determines the importance of a node by measuring the number of edges it is connected to. That is, the node with the most connected nodes is usually considered more important. In a directed graph, degree centrality is divided into in-degree and out-degree.

[0143] Closeness Centrality: Closeness Centrality measures the average distance from a node to any other node. The closer a node is to all other nodes, the higher its closeness centrality. This can reflect the reachability of the node in the network.

[0144] Betweenness Centrality: Betweenness Centrality measures the number of times a node appears in all the shortest paths in a graph, that is, the number of times a node serves as a bridge for the shortest paths between other two nodes. The higher the value, the more important the role of the node in the information flow.

[0145] Eigenvector Centrality: Eigenvector Centrality believes that the importance of a node depends on the importance of its neighbor nodes. It is calculated iteratively until the eigenvector converges to a stable solution. This centrality takes into account the influence of the node and the degree of connection to important nodes.

[0146] PageRank: The PageRank algorithm is based on an assumption that the importance of a node is determined by the number and quality of the nodes pointing to it. It uses an iterative method to calculate the PageRank value of each node until convergence.

[0147] K-Core Algorithm: The K-Core algorithm is used to identify highly interconnected subgraphs in a graph, that is, the k-core. In the k-core, each node is connected to at least k other nodes. The k-core algorithm can identify the core structure in the network.

[0148] HITS Algorithm: This algorithm distinguishes two types of nodes: Hubs and Authorities. Hubs are nodes that point to multiple Authorities, while Authorities are nodes that are pointed to by multiple Hubs. The HITS algorithm calculates the Hubs and Authorities scores of each node iteratively.

[0149] Please refer to Figure 1 , the embodiments of the present application provide a periodical matching recommendation method based on big data, including:

[0150] Step S1, obtain the manuscript to be detected, and extract the feature information therein to form a target feature vector.

[0151] Among them, the manuscript to be detected is a received document that needs to be matched with relevant periodicals, and the content therein can be in any form, not strictly required to be in the format of a thesis or a periodical, as long as there is text.

[0152] Step S2, obtain historical periodical data, and extract the feature information of each periodical to form a corresponding periodical feature vector.

[0153] Specifically, since the journals published in the journal database are already in a standardized fixed format, when extracting journal feature information, the technical field to which the journal belongs, the impact factor, etc. can be directly used as feature information to construct a journal feature vector, or the same feature extraction and construction method as the above-mentioned target feature vector can be adopted.

[0154] Step S3: Train a machine learning model based on each journal feature vector to obtain a journal matching model.

[0155] Among them, the machine learning model can adopt models such as decision trees, support vector machines, K-nearest neighbors, and clustering algorithms.

[0156] Specifically, in actual applications, due to different training methods, different machine learning models have different performances in classification and clustering. For example, a decision tree infers the target value from data features by learning simple decision rules, so the training method is to construct a tree structure, usually using greedy algorithms such as ID3, C4.5, or CART; a support vector machine finds the optimal boundary between data points, and the training method is to use the Lagrange multiplier method and solve the dual problem; K-nearest neighbor is instance-based learning, predicting by finding the K nearest neighbors of the test data point, and the training method is to store the training data and calculate distances and perform searches during prediction; clustering algorithms such as K-Means and hierarchical clustering group data points into different clusters; the training method is iterative optimization, such as K-Means iteratively selects cluster centers and reassigns data points. Therefore, multiple machine learning models can be trained using journal feature vectors, and the model with the best performance can be selected as the journal matching model based on the training results.

[0157] During the training process, the journal feature vector is input into the machine learning model, and the output data is the matching journal and the matching degree of each matching journal. Each matching journal is a journal associated with the journal corresponding to the input journal feature vector, and the matching degree is the degree of association between the matching journal and the journal corresponding to the input journal feature vector. Therefore, the training objective is that the matching journal with the highest matching degree output by the model is the journal corresponding to the input journal feature vector. At this time, it can be considered that the model has achieved the expectation.

[0158] Step S4: Input the target feature vector into the journal matching model to obtain multiple matching journals and corresponding matching degrees.

[0159] Specifically, the number of output matching journals is usually determined by a preset matching degree threshold inside the journal matching model, that is, it is considered that the journal is not relevant to the manuscript to be detected after the matching degree is lower than the preset matching degree threshold.

[0160] Furthermore, a preset number of matching journals can also be set inside the journal matching model. The matching journals are sorted according to the matching degree, and the top preset number of matching journals are selected for output. However, although this method can further reduce the amount of similarity calculation in the next step, i.e., step S5, in the case of a large number of researchers and journals in the technical field, useful journal data may be missed, so the usage rate is relatively low.

[0161] Step S5, calculate the similarity between the target feature vector and the journal feature vectors of each matching journal.

[0162] Among them, the methods for calculating the similarity between two feature vectors can include: 1) Cosine similarity, which measures the similarity between two vectors by calculating the cosine value of the angle between them; 2) Euclidean distance, which calculates the straight-line distance between two vectors in Euclidean space; 3) Jaccard similarity coefficient, which is used to measure the similarity between two sets and calculates the ratio of the intersection and the union; 4) Pearson correlation coefficient, which measures the linear correlation between two variables; 5) Mahalanobis distance, which depends on the inverse of the covariance matrix and is an effective distance measurement method; 6) KL divergence, which is used to measure the dissimilarity between two probability distributions.

[0163] Specifically, in step S5 of the present application, the object of similarity calculation is only the matching journals output by the journal matching model, rather than all historical journals in the database. Therefore, compared with the prior art, the amount of calculation is greatly reduced.

[0164] Step S6, sort each matching journal based on the similarity and the matching degree to obtain a recommended journal list.

[0165] Specifically, the similarity is directly calculated from two feature vectors, and the matching degree is output by the model. The present application combines the parameters obtained by these two methods to sort the matching journals, which will make the sorting result more accurate and reduce the deviation of the matching result caused by inaccurate calculation of any one of the parameters.

[0166] A journal matching and recommendation method based on big data provided by the above embodiments first trains a journal matching model through journal feature vectors to obtain multiple matching journals and corresponding matching degrees, then calculates the similarity between the target feature vector of the manuscript to be detected and the journal feature vectors of each matching journal, and finally forms a recommended journal list by combining the similarity and the matching degree. The above method combines the similarity and the matching method of the large model based on feature vectors, which not only avoids the situation that the model output is affected by inaccurate manually set labels, greatly improves the accuracy of journal matching, but also reduces the amount of data required for similarity calculation, that is, the manuscript to be detected only needs to calculate the similarity with the matching journals output by the model, thus improving the matching speed.

[0167] In some embodiments, the method further includes:

[0168] Step S01, before extracting the feature information of the manuscript to be detected, removing the preset stop words in the manuscript to be detected.

[0169] Among them, the preset stop words refer to words that frequently appear in the text but contribute little to the meaning of the text, such as "de", "he", "shi", etc. Removing these words can reduce the noise when extracting feature information subsequently.

[0170] Step S02, performing word segmentation, stemming, and lowercase conversion on the manuscript to be detected.

[0171] Among them, word segmentation refers to splitting the text in the manuscript to be detected into individual words or phrases; stemming refers to converting a word into its dictionary form, for example, converting "better" to the comparative form of "good"; lowercase conversion is to convert all text to lowercase to eliminate the morphological differences caused by case.

[0172] In addition to the above-mentioned processing methods, operations such as removing phrases and synonyms, annotating the parts of speech such as nouns, verbs, and adjectives, and removing punctuation marks and special symbols can also be added.

[0173] In the specific implementation process, the content of the manuscript to be detected may contain a lot of information that is useless for journal matching or information that affects the matching accuracy. Therefore, before extracting the feature information in this application, performing the above-mentioned processing operations on the manuscript to be detected can make the extraction of feature information faster and more accurate, thereby making the matching result more accurate.

[0174] Specifically, the word segmentation operation helps to identify the basic structure of the text, removing stop words can reduce data noise, improve the quality of features and the accuracy of analysis, stemming can retain the part-of-speech information of words, which helps to understand the meaning of the manuscript to be detected, lowercase conversion can eliminate case differences, ensure the consistency of text processing, and avoid treating the same word as different words due to different cases. Removing special characters and punctuation can reduce the interference of irrelevant information and simplify the text data.

[0175] Please refer to Figure 2 , in some embodiments, obtaining the manuscript to be detected and extracting the feature information therein to form a target feature vector may specifically include the following steps:

[0176] Step S11, extracting the sentiment feature vector of the manuscript to be detected based on a preset sentiment dictionary.

[0177] Among them, the sentiment feature vector is used to represent the understanding of the emotional tendency in the manuscript to be detected, which helps the journal matching model to more accurately capture the views and attitudes of the author of the manuscript to be detected.

[0178] Specifically, in a certain technical field, there may be different views and opinions regarding a certain research direction; for example, regarding the application of a certain material in a certain product, Journal A may focus on the positive effects or advantages of the material in the product, while Journal B may focus on the negative effects or disadvantages of the material in the product.

[0179] Therefore, in the absence of emotional features, important semantic information may be lost during the matching process, affecting the model's comprehensive understanding of the document to be detected, thereby reducing the accuracy of the matching.

[0180] Step S12: Extract the term frequency feature vector of the document to be detected using the TF-IDF method.

[0181] Specifically, extract all unique words from the document to be detected and create a vocabulary. Calculate the number of occurrences of each word in the vocabulary, which is the term frequency; for each word in the vocabulary, calculate its inverse document frequency in the document to be detected.

[0182] For each word in each document, calculate its TF-IDF value, which is the product of the term frequency and the inverse document frequency, thereby obtaining the term frequency feature vector, where each dimension in the term frequency feature vector corresponds to a word in the vocabulary.

[0183] Step S13: Combine the emotional feature vector and the term frequency feature vector to obtain the target feature vector.

[0184] In the above embodiments, when constructing the target feature vector, the emotional feature vector is added and combined with the term frequency feature vector, enabling the journal matching model to provide more accurate matching and recommendations by analyzing the emotional tendency of the document to be detected.

[0185] Please refer to Figure 3 , in some embodiments, the above-mentioned extraction of the emotional features of the document to be detected based on the preset emotional dictionary includes:

[0186] Step S111: Assign emotional polarity scores to each word in the document to be detected according to the preset emotional dictionary.

[0187] Among them, existing emotional dictionaries such as AFINN, VADER, etc. can be used here. According to the emotional scores of different words in the dictionary, emotional polarity scores are assigned to each word in the document to be detected.

[0188] Specifically, when the polarity is negative, the corresponding emotional polarity score is negative; when the polarity is positive, the corresponding emotional polarity score is positive; when the polarity is neutral, the corresponding emotional polarity score is 0.

[0189] Step S112: Calculate the text sentiment score of the text to be detected based on each sentiment polarity score.

[0190] Specifically, weight the sentiment polarity scores according to the word frequencies of each word, or directly average each sentiment polarity score to obtain the text sentiment score.

[0191] Step S113: Determine the sentiment type and sentiment intensity of the text to be detected according to the text sentiment score.

[0192] Among them, the sentiment type is divided into categories such as positive, negative or neutral, which is usually obtained according to the numerical interval where the text sentiment score is located. For example, when the sentiment score is less than a preset negative threshold, it is a negative sentiment; when it is greater than a certain preset positive threshold, it is a positive sentiment; when it is between the preset negative threshold and the preset positive threshold, it is a neutral sentiment. The sentiment intensity is the difference between the text sentiment score and the corresponding threshold. When it is a negative sentiment, it is the difference from the preset negative threshold; when it is a positive sentiment, it is the difference from the preset positive threshold; when it is a neutral sentiment, it is the difference from 0. The difference is used as the sentiment intensity.

[0193] Step S114: Construct a sentiment feature vector from the text sentiment score, sentiment type and sentiment intensity.

[0194] The above embodiments construct a sentiment feature vector from the text sentiment score, sentiment type and sentiment intensity, ensuring the comprehensiveness and accuracy of the sentiment analysis of the text to be detected, and helping the journal matching model to match more precisely.

[0195] Please refer to Figure 4 , in some embodiments, the method further includes:

[0196] Step S301: Calculate the chi-square statistic of each journal feature vector before training the machine learning model.

[0197] Specifically, the calculation formula of the chi-square statistic x 2 is:

[0198]

[0199] Among them, O ij is the word frequency of the coordinate (i, j) in the journal feature vector, and E ij is the expected frequency under the corresponding independent hypothesis.

[0200] Step S302: Construct a contingency table corresponding to each journal feature vector; among them, the rows of the contingency table represent the values of different features in the journal feature vector, and the columns represent the categories of different features.

[0201] Specifically, a contingency table is usually a two-dimensional table, and the cells of the table contain the observed frequencies or counts for the corresponding category combinations. For example, assume there are two categorical variables in the journal feature vector: gender (male and female) and smoking habit (smoker and non-smoker). A simple contingency table might be as follows:

[0202]

[0203] In this table, a represents the number of male smokers, b represents the number of male non-smokers, c represents the number of female smokers, and d represents the number of female non-smokers. Through this contingency table, a chi-square test can be used to test whether there is a statistically significant association between gender and smoking habit. Therefore, in this application, the chi-square test is used to determine whether there is an association among the various feature information in the journal feature vector, so as to be able to judge the quality of the journal corresponding to the journal feature vector.

[0204] Step S303: Determine the chi-square threshold of the journal feature vector according to the preset significance level and the contingency table.

[0205] Among them, the preset significance level can usually be taken as 0.05, which is used to judge whether the result of the chi-square test is significant.

[0206] Step S304: If the chi-square statistic is less than the chi-square threshold, then eliminate the corresponding journal feature vector.

[0207] In practical applications, since the quality of the journals in the journal database is uneven, the quality of the corresponding generated journal feature vectors is also uneven. Therefore, in order to improve the training speed of the model and avoid some journal feature vectors with poor quality and poor effectiveness that can neither produce good training effects nor waste training time, this application uses the chi-square test method to screen out the journal feature vectors with better effectiveness and ensure the training efficiency of the machine learning model.

[0208] In some embodiments, the above-mentioned determination of the chi-square threshold of the journal feature vector according to the preset significance level and the contingency table includes:

[0209] Step S3031: Subtract 1 from the number of rows and the number of columns of the contingency table respectively, and then multiply them to obtain the degrees of freedom.

[0210] Step S3032: Multiply the preset significance level by the degrees of freedom to obtain the chi-square threshold.

[0211] Specifically, the chi-square threshold = preset significance level × (number of rows - 1) × (number of columns - 1).

[0212] The chi-square test is a fast and effective method. Especially when dealing with high-dimensional data sets, it can significantly reduce the number of invalid or low-quality features, thereby accelerating the training efficiency of the model.

[0213] In some embodiments, calculating the similarity between the target feature vector and the journal feature vectors of each matching journal includes:

[0214] Calculating the Manhattan distance or Jaccard similarity coefficient between the target feature vector and the journal feature vector, and using the Manhattan distance or Jaccard similarity coefficient as the similarity.

[0215] Among them, the Manhattan distance, also known as the L1 distance, is the sum of the absolute values of the differences in each dimension. The Jaccard similarity coefficient is calculated by taking the absolute value of the ratio of the intersection to the union of two feature vectors.

[0216] Specifically, among many similarity calculation methods, the present application focuses on selecting the calculation method of the Manhattan distance or the calculation method of the Jaccard similarity coefficient. This is because the calculation speeds of the Manhattan distance and the Jaccard similarity coefficient are fast, they are robust to outliers, and are suitable for large-scale data sets, especially when the feature dimension is very high; therefore, adopting the above two methods helps to further improve the calculation speed of the similarity and further improve the speed of journal recommendation.

[0217] Please refer to Figure 5 , in some embodiments, sorting each matching journal based on the similarity and matching degree to obtain a recommended journal list may specifically include the following steps:

[0218] Step S61, normalizing each similarity and each matching degree respectively.

[0219] Step S62, adding the normalized similarity and matching degree to obtain the sorting value of the matching journal.

[0220] Step S63, sorting each matching journal according to the sorting value to obtain a recommended journal list.

[0221] Specifically, since the similarity is directly calculated from two feature vectors, and the matching degree is output by the journal matching model, the formats of the matching degrees output by different journal matching models may be different. Therefore, before combining the two parameters, the present application normalizes the two parameters respectively to ensure the accuracy of the sorting value and the accuracy of the recommended journal list.

[0222] In some embodiments, the method further includes:

[0223] Step S71, calculating the popularity score of each matching journal based on the knowledge graph method to obtain a popularity sorted list.

[0224] Step S72, adjusting the sorting of each matching journal in the recommended journal list according to the popularity sorted list.

[0225] Specifically, when recommending journals based on journal similarity and matching degree, it is easy to misplace some basic authoritative journals with earlier years in the corresponding technical field in relatively front positions; in academic research, the latest progress in a certain research direction within a certain technical field is often the part that research scholars in the corresponding research direction focus on.

[0226] Therefore, the present application further takes the popularity of journals into account in the ranking of the journal recommendation list, so that when recommending journals with a relatively high degree of association, it is possible to recommend the research content with a relatively high degree of attention and popularity in the corresponding technical field in the front.

[0227] Please refer to Figure 6 , in some embodiments, the above method based on the knowledge graph calculates the popularity scores of each matching journal to obtain a popularity ranking list, which may specifically include the following steps:

[0228] Step S711: Take each matching journal as a node, take the journal feature vector as the attribute of the corresponding node, and take the citation relationship between each matching journal as the edge connecting each node to obtain a journal knowledge graph.

[0229] Specifically, considering that each matching journal is very likely to belong to the same technical field or even the same research direction, and there is very likely a mutual citation relationship between them. Therefore, the present application uses the knowledge graph method to associate the matching journals with citation relationships to ensure the accuracy of the calculation of the popularity scores of the matching journals.

[0230] Step S712: Obtain the comment popularity score according to the comment text of each node and the corresponding comment time.

[0231] Specifically, according to the number of newly added comment texts under the matching journal corresponding to the node and the corresponding comment time on the same day, construct a comment curve, and take the derivative of the comment curve to obtain the comment popularity score.

[0232] Step S713: Obtain the collection popularity score according to the cumulative collection volume and collection growth rate of each node.

[0233] Step S714: Perform weighted calculation on the comment popularity score and the collection popularity score to obtain the node score.

[0234] Step S715: Use a graph algorithm to calculate the popularity score of each node according to the node scores and the edges between each node.

[0235] Specifically, graph algorithms (such as PageRank, HITS, Betweenness Centrality, etc.) can be used to calculate the importance of each node. In the present application, "importance" is equated with the popularity score of the matching journal here.

[0236] Step S716: Sort the matching journals corresponding to each node according to the popularity score to obtain a popularity sorted list.

[0237] Based on the nodes in the above embodiments, that is, the review situation and collection situation of the matching journals, the node scores of each node are determined, and the popularity scores of the nodes are further calculated, which can ensure the calculation accuracy and fit degree of the popularity scores.

[0238] Please refer to Figure 7 , in some embodiments, the above-mentioned obtaining the collection popularity score according to the cumulative collection volume and collection growth rate of each node may specifically include the following steps:

[0239] Step S7141: Determine the first score of the collection interval where the cumulative collection volume is located and the second score of the growth rate interval where the collection growth rate is located. The collection growth rate is the number of newly added collections within a preset time period before the current moment divided by the preset time period.

[0240] Specifically, this application sets different numerical intervals for different cumulative collection volumes and collection growth rates, and each numerical interval corresponds to a different score. The two scores respectively reflect the authority and attention of the matching journals.

[0241] Step S7142: Add the first score and the second score to obtain the collection popularity score.

[0242] In some embodiments, the above-mentioned adjusting the sorting of each matching journal in the recommended journal list according to the popularity sorted list includes:

[0243] Take the average value of the rankings of the matching journals in the popularity sorted list and the recommended journal list as the new ranking of the matching journals in the recommended journal list. Specifically, for example, if the ranking of a certain matching journal in the popularity sorted list is 1 and the ranking in the recommended journal list is 9, the ranking is adjusted to 5 after taking the average value; if the rankings of two matching journals are the same after averaging, they can be randomly sorted or the user preferences can be determined according to the historical browsing data of the user account, such as impact factor preference, geographical location preference, etc. Among the two matching journals with the same ranking, the journal that better meets the user preference is ranked in a more forward position.

[0244] Please refer to Figure 8 , in some embodiments, the method further includes:

[0245] Step S81: Obtain the click times of each matching journal and modify the corresponding matching degree based on the click times.

[0246] Step S82: Train the journal matching model based on the modified matching degree and the corresponding journal feature vector.

[0247] Specifically, after presenting the recommended journal list, the number of clicks of each matching journal by the user is obtained, and the number of clicks is added to the matching degree to obtain a new matching degree. Then, the journal matching model is feedback-trained based on the new matching degree, so as to perform self-optimization according to the user feedback and continuously iterate the model parameters to improve the matching accuracy.

[0248] Please refer to Figure 9 , Another embodiment of the present application provides a journal matching and recommendation device based on big data, including:

[0249] A target manuscript module 101, configured to obtain a manuscript to be detected and extract feature information therein to form a target feature vector.

[0250] A journal module 102, configured to obtain historical journal data and extract feature information of each journal to form a corresponding journal feature vector.

[0251] A training module 103, configured to train a machine learning model based on each journal feature vector to obtain a journal matching model.

[0252] A matching module 104, configured to input the target feature vector into the journal matching model to obtain multiple matching journals and corresponding matching degrees.

[0253] A calculation module 105, configured to calculate the similarity between the target feature vector and the journal feature vectors of each matching journal.

[0254] A recommendation module 106, configured to sort each matching journal based on the similarity and the matching degree to obtain a recommended journal list.

[0255] The journal matching and recommendation device provided in the above embodiment trains a journal matching model through the training module 103 to obtain multiple matching journals and corresponding matching degrees, then calculates the similarity between the target feature vector of the manuscript to be detected and the journal feature vectors of each matching journal through the calculation module 105, and finally forms a recommended journal list by combining the similarity and the matching degree. The above device combines the similarity and the matching method of the large model based on the feature vector, which not only avoids the situation that the inaccurate labels set manually affect the model output, greatly improves the accuracy of journal matching, but also reduces the amount of data required for similarity calculation, that is, the manuscript to be detected only needs to calculate the similarity with the matching journals output by the model, thus improving the matching speed.

[0256] Further, the device further includes a preprocessing module, configured to remove preset stop words in the manuscript to be detected before extracting the feature information thereof; and perform word segmentation, stemming, and lowercase conversion on the manuscript to be detected.

[0257] Please refer to Figure 10 , Further, the above target manuscript module 101 includes:

[0258] An emotion unit 11 for extracting an emotion feature vector of a manuscript to be detected based on a preset emotion dictionary.

[0259] A term frequency unit 12 for extracting a term frequency feature vector of a manuscript to be detected by using the TF-IDF method.

[0260] A combination unit 13 for combining the emotion feature vector and the term frequency feature vector to obtain a target feature vector.

[0261] Further, the above-mentioned emotion unit 11 is used for:

[0262] Assigning an emotion polarity score to each word in the manuscript to be detected according to the preset emotion dictionary.

[0263] Calculating a manuscript emotion score of the manuscript to be detected based on each emotion polarity score.

[0264] Determining the emotion type and emotion intensity of the manuscript to be detected according to the manuscript emotion score.

[0265] Constructing the manuscript emotion score, emotion type and emotion intensity into an emotion feature vector.

[0266] Further, the device further includes:

[0267] A statistical module for calculating the chi-square statistic of each journal feature vector before training a machine learning model.

[0268] A contingency table module for constructing a contingency table corresponding to each journal feature vector.

[0269] A threshold module for determining a chi-square threshold of the journal feature vector according to a preset significance level and the contingency table.

[0270] An elimination module for eliminating the corresponding journal feature vector if the chi-square statistic is less than the chi-square threshold.

[0271] Further, the above-mentioned threshold module includes:

[0272] A degree of freedom unit for subtracting 1 from the number of rows and the number of columns of the contingency table respectively and then multiplying them to obtain the degree of freedom.

[0273] A threshold unit for multiplying the preset significance level and the degree of freedom to obtain the chi-square threshold.

[0274] Further, the above-mentioned calculation module 105 is used for calculating the Manhattan distance or Jaccard similarity coefficient between the target feature vector and the journal feature vector, and using the Manhattan distance or Jaccard similarity coefficient as the similarity.

[0275] Please refer to Figure 11, Further, the above-mentioned recommendation module 106 includes:

[0276] A normalization unit 61, configured to normalize each similarity and each matching degree respectively.

[0277] A sorting value unit 62, configured to add the normalized similarity and matching degree to obtain a sorting value of the matching journals.

[0278] A recommendation list unit 63, configured to sort each matching journal according to the sorting value to obtain a recommended journal list.

[0279] Further, the device further includes:

[0280] A popularity sorting module, configured to calculate a popularity score of each matching journal based on a knowledge graph method to obtain a popularity sorted list.

[0281] A sorting update module, configured to adjust the sorting of each matching journal in the recommended journal list according to the popularity sorted list.

[0282] Further, the above-mentioned popularity sorting module includes:

[0283] A graph construction unit, configured to use each matching journal as a node, use the journal feature vector as the attribute of the corresponding node, and use the citation relationship between each matching journal as the edge connecting each node to obtain a journal knowledge graph.

[0284] A comment popularity unit, configured to obtain a comment popularity score according to the comment text of each node and the corresponding comment time.

[0285] A favorite popularity unit, configured to obtain a favorite popularity score according to the cumulative favorite quantity and the favorite growth rate of each node.

[0286] A node score unit, configured to perform weighted calculation on the comment popularity score and the favorite popularity score to obtain a node score.

[0287] A node popularity unit, configured to calculate the popularity score of each node according to each node score and the edges between each node by using a graph algorithm.

[0288] A popularity sorting unit, configured to sort the matching journals corresponding to each node according to the popularity score to obtain a popularity sorted list.

[0289] Further, the above-mentioned favorite popularity unit is used for:

[0290] Determine a first score in the favorite interval where the cumulative favorite quantity is located, and a second score in the growth rate interval where the favorite growth rate is located.

[0291] Add the first score and the second score to obtain a favorite popularity score.

[0292] Further, the above sorting update module is used to take the average of the sorting of the matching journals in the popularity sorting list and the recommended journal list as the new sorting of the matching journals in the recommended journal list.

[0293] Further, the device further includes:

[0294] A click acquisition module, configured to acquire the click times of each matching journal and modify the corresponding matching degree based on the click times.

[0295] A feedback training module, configured to train the journal matching model based on the modified matching degree and the corresponding journal feature vector.

[0296] For the specific limitations of the journal matching and recommending device based on big data provided in this embodiment, reference may be made to the embodiment of the journal matching and recommending method based on big data in the foregoing, which will not be elaborated herein. Each module in the above-mentioned journal matching and recommending device based on big data can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0297] In the embodiments disclosed in this application, it should be understood that the disclosed products can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the modules can be electrical, mechanical or other forms. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place, or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing module, or each module exists physically alone, or two or more modules can be integrated into one module.

[0298] An embodiment of the present application provides a computer device, which may include a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the processor is caused to execute the steps of a big data-based journal matching and recommendation method according to any of the above embodiments.

[0299] For the working process, working details, and technical effects of the computer device provided in this embodiment, reference may be made to the embodiments of a big data-based journal matching and recommendation method in the foregoing text, and details are not described herein again.

[0300] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a big data-based journal matching and recommendation method according to any of the above embodiments are implemented. Among them, the computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash drive, and / or a Memory Stick, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0301] For the working process, working details, and technical effects of the computer-readable storage medium provided in this embodiment, reference may be made to the embodiments of a big data-based journal matching and recommendation method in the foregoing text, and details are not described herein again.

[0302] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM).

[0303] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0304] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A journal matching and recommendation method based on big data, characterized in that, Including: Obtain the manuscript to be detected, and extract the feature information therein to form a target feature vector; Obtain historical journal data, and extract the feature information of each journal to form a corresponding journal feature vector; Train a machine learning model based on each of the journal feature vectors to obtain a journal matching model; Input the target feature vector into the journal matching model to obtain multiple matching journals and corresponding matching degrees; Calculate the similarity between the target feature vector and the journal feature vectors of each of the matching journals; Rank each of the matching journals based on the similarity and the matching degree to obtain a recommended journal list; Calculate the popularity scores of each of the matching journals based on the method of knowledge graph to obtain a popularity ranking list; Specifically, take each of the matching journals as nodes, take the journal feature vector as the attribute corresponding to the node, take the citation relationship between each of the matching journals as the edge connecting each of the nodes to obtain a journal knowledge graph; obtain the comment popularity score according to the review text and the corresponding review time of each of the nodes; obtain the collection popularity score according to the cumulative collection volume and the collection growth rate of each of the nodes; Perform weighted calculation on the comment popularity score and the collection popularity score to obtain a node score; Use a graph algorithm to calculate the popularity scores of each of the nodes according to each of the node scores and the edges between each of the nodes; rank the matching journals corresponding to each of the nodes according to the popularity scores to obtain the popularity ranking list; Adjust the ranking of each of the matching journals in the recommended journal list according to the popularity ranking list.

2. The method for periodical matching and recommendation based on big data according to claim 1, wherein The obtaining the manuscript to be detected and extracting the feature information therein to form a target feature vector includes: Extract the sentiment feature vector of the manuscript to be detected based on a preset sentiment dictionary; Use the TF-IDF method to extract the word frequency feature vector of the manuscript to be detected; Combine the sentiment feature vector and the word frequency feature vector to obtain the target feature vector.

3. The method for periodical matching and recommendation based on big data according to claim 2, wherein The extracting the sentiment feature vector of the manuscript to be detected based on a preset sentiment dictionary includes: Assign sentiment polarity scores to each word in the manuscript to be detected according to the preset sentiment dictionary; Calculate the manuscript sentiment score of the manuscript to be detected based on each of the sentiment polarity scores; Determine the sentiment type and sentiment intensity of the manuscript to be detected according to the manuscript sentiment score; Construct the manuscript sentiment score, the sentiment type and the sentiment intensity into the sentiment feature vector.

4. The method for periodical matching recommendation based on big data according to claim 1, wherein Also including: Before training the machine learning model, calculate the chi-square statistic of each of the journal feature vectors; Construct a contingency table corresponding to each of the journal feature vectors; wherein, the rows of the contingency table represent the values of different features in the journal feature vector, and the columns represent different feature categories; Determine the chi-square threshold of the journal feature vector according to a preset significance level and the contingency table; If the chi-square statistic is less than the chi-square threshold, then eliminate the corresponding journal feature vector.

5. The method for periodical matching recommendation based on big data according to claim 1, wherein The obtaining the collection popularity score according to the cumulative collection volume and the collection growth rate of each of the nodes includes: Determine the first score of the collection interval where the cumulative collection volume is located, and the second score of the growth rate interval where the collection growth rate is located; Add the first score and the second score to obtain the collection popularity score.

6. The method for periodical matching and recommendation based on big data according to claim 1, wherein Adjusting the sorting of each of the matching journals in the recommended journal list according to the popularity sorted list includes: Taking the average of the sorting of the matching journal in the popularity sorted list and the recommended journal list as the new sorting of the matching journal in the recommended journal list.

7. A periodical matching and recommending device based on big data, characterized in that Including: A target manuscript module for obtaining a manuscript to be detected and extracting feature information therein to form a target feature vector; A journal module for obtaining historical journal data and extracting feature information of each journal to form a corresponding journal feature vector; A training module for training a machine learning model based on each of the journal feature vectors to obtain a journal matching model; A matching module for inputting the target feature vector into the journal matching model to obtain a plurality of matching journals and corresponding matching degrees; A calculation module for calculating the similarity between the target feature vector and the journal feature vectors of each of the matching journals; A recommendation module for sorting each of the matching journals based on the similarity and the matching degree to obtain a recommended journal list; A popularity sorting module for calculating the popularity score of each matching journal based on the method of the knowledge graph to obtain a popularity sorted list; Specifically, the popularity sorting module includes: a graph construction unit for taking each matching journal as a node, taking the journal feature vector as the attribute of the corresponding node, and taking the citation relationship between each matching journal as the edge connecting each node to obtain a journal knowledge graph; a comment popularity unit for obtaining a comment popularity score according to the comment text of each node and the corresponding comment time; a collection popularity unit for obtaining a collection popularity score according to the cumulative collection volume and the collection growth rate of each node; a node score unit for performing weighted calculation on the comment popularity score and the collection popularity score to obtain a node score; a node popularity unit for calculating the popularity score of each node according to each node score and the edges between each node by using a graph algorithm; a popularity sorting unit for sorting the matching journals corresponding to each node according to the popularity score to obtain a popularity sorted list; A sorting update module for adjusting the sorting of each matching journal in the recommended journal list according to the popularity sorted list.

8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the big data-based journal matching and recommendation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Paper management method, server and system based on big data

    CN111353031A

  • Display device, voice search method and storage medium

    CN115862615A

  • Knowledge question-answering method and device, equipment and storage medium

    CN117573821A