A method and system for analyzing the correlation of scientific and technological information
By using semantic feature extraction and convolutional neural network to calculate the comprehensive correlation degree in scientific and technological intelligence correlation analysis, the problem of low efficiency and accuracy of literature screening in the existing technology is solved, and more efficient and accurate literature correlation analysis is achieved.
Patent Information
- Application Number
- CN202410871032.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-07-01
AI Technical Summary
The existing technology is difficult to meet the requirements of accuracy and efficiency when efficiently screening relevant information in massive scientific and technological literature. Traditional methods lack a deep understanding of the literature content and semantic correlation recognition, resulting in inaccurate sorting of results.
A scientific and technological intelligence correlation analysis method is adopted to preprocess the technology literature or query statement uploaded by the user, extract semantic feature vectors, calculate the comprehensive correlation degree using the second convolutional neural network algorithm, and sort the literature through the improved Sigmoid activation function to consider the time and popularity.
It improves the accuracy and efficiency of literature screening, enhances user experience, can understand and represent the semantic information of the literature more accurately, and comprehensively considers multi-dimensional factors, which improves the effectiveness of literature correlation analysis.
Smart Images

Figure CN118782165B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of scientific and technological information retrieval and acquisition, and in particular to a scientific and technological information correlation analysis method and system. Background Art
[0002] In today's information age, the pace of scientific research and technological innovation is rapid, and the number and complexity of documents are increasing. How to effectively mine relevant information from massive scientific and technological documents and then understand and utilize this information has become an important challenge for scientific research and industry. Traditional document retrieval and association analysis methods usually rely on keyword matching and simple similarity calculations. These methods are often difficult to meet the requirements of efficiency and accuracy when processing massive and complex text data. Traditional document retrieval methods mainly rely on keyword-based search engines. The core of these methods is to find documents through keyword matching and Boolean logic. Although this method met the basic retrieval needs in the early days, its limitations have gradually emerged: first, it is keyword-dependent. The relevant information in the document may appear in various forms. Simply relying on keyword matching cannot capture the semantics and contextual relationships of the document; second, it lacks contextual understanding. Traditional methods lack a deep understanding of the content of the document and cannot accurately identify the semantic associations between documents; the best result is inaccurate result sorting. Simple keyword matching and similarity calculations are difficult to accurately sort the results, causing users to spend a lot of time screening results.
[0003] With the rapid development of natural language processing (NLP) technology, especially the introduction of deep learning models (such as BERT and GPT series), the methods of document analysis have also been significantly improved. These models can be pre-trained on large-scale text corpora to capture rich semantic and contextual information, significantly improving the ability to understand and process text. Among them, the BERT model performs well in capturing complex relationships and contextual dependencies between words through a bidirectional Transformer structure. BERT's global feature representation enables it to more accurately understand and represent semantic information in documents. Convolutional neural networks (CNNs) have not only achieved success in image processing, but have also been widely used in text processing. By applying one-dimensional convolution and pooling operations, CNNs can effectively capture local patterns and features in text. In the process of document analysis, users pay attention not only to the semantic content of the document, but also to the timeliness and influence of the document. For example, the latest published document may contain the latest research results, while documents with high views and searches often have higher attention and reference value. Therefore, in the analysis of document relevance, it is crucial to comprehensively consider the time information and popularity (such as search volume and views) of the document. In order to more accurately calculate the comprehensive correlation between documents, a model that can comprehensively consider multi-dimensional factors is needed. Traditional similarity calculation methods (such as cosine similarity) can only handle simple feature similarities and cannot fully reflect the complex relationships between documents. Modern deep learning models, such as convolutional neural networks, can better capture and express the deep correlation between documents through complex nonlinear transformations and multi-level feature extraction.
[0004] However, the existing deep learning model directly processes massive documents, resulting in a significant reduction in the efficiency of document screening; and the existing document screening methods only use similarity comparison or use deep learning models alone, resulting in inaccurate judgments and reduced efficiency. For example, there is no algorithm in the existing technology that performs rough screening and then fine screening on massive data to achieve first full and then accurate, thereby greatly improving the efficiency and accuracy of document screening; and the current scientific and technological literature screening process does not take into account the time and the number of document readings and views, resulting in inaccurate screening results. The release time of existing technical documents is crucial for understanding research trends and obtaining the latest information. The traditional time-weighted method has limitations in processing the timeliness of documents. It does not add time information to the deep model activation function and cannot dynamically adjust the weight of the document; and the search volume and views of the document reflect its attention in academia and industry. The existing technology does not use these indicators as an important reference for the influence of the document, resulting in a significant reduction in the accuracy of document screening. In view of the limitations of the existing scientific and technological literature screening judgment, a new solution is urgently needed for the automatic screening of scientific and technological documents with high accuracy and efficiency to improve the processing efficiency and accuracy. Summary of the invention
[0005] In view of the above problems mentioned in the prior art, the present invention provides a method and system for analyzing the relevance of scientific and technological intelligence. The method first receives the scientific and technological literature or query statements uploaded by the user, and performs preprocessing to obtain the preprocessed text; secondly, the semantic features of the preprocessed text are extracted to obtain the final feature vector F; then the related documents are determined from the database according to the final feature vector F; the second convolutional neural network algorithm is used again to calculate the comprehensive relevance between the scientific and technological literature or query statements uploaded by the user and the determined related documents; finally, the documents are displayed to the user according to the comprehensive relevance sorting. The present application determines the related documents according to the final feature vector F, and then uses the second convolutional neural network algorithm to calculate the comprehensive relevance of the related documents, and then displays the documents to the user. By adopting the improved Sigmoid activation function to consider the monthly search volume and pageview volume, the accuracy and efficiency of document screening are greatly improved, and the user experience is greatly increased.
[0006] The present application provides a method for analyzing the correlation of scientific and technological intelligence, comprising the steps of:
[0007] S1: receiving scientific documents or query statements uploaded by users, and preprocessing the uploaded scientific documents or query statements to obtain preprocessed texts;
[0008] S2: Extract semantic features from the preprocessed text to obtain the final feature vector F;
[0009] S3: Determine the relevant documents and their monthly search volume and page views from the database according to the final feature vector F;
[0010] S4: using the second convolutional neural network algorithm to calculate the comprehensive correlation between the scientific literature or query statement uploaded by the user and the determined related literature;
[0011] S41: Pair the final feature vector F with the feature vector of each associated document in the database as an input pair, wherein the input pair includes the monthly search volume and pageview volume of the associated document;
[0012] S42: All input pairs are input into the second convolutional neural network model. The second convolutional neural network model uses an improved Sigmoid activation function to perform a convolution operation, performs a maximum pooling operation on the output after convolution, and combines the pooling results of all convolution kernels through a fully connected layer to calculate a comprehensive correlation score.
[0013] The improved Sigmoid activation function is as follows:
[0014]
[0015] Among them, f(x) is the output of the Sigmoid activation function, x is the output of the convolution operation, and w(t i ) is the time decay weight factor, t i is the publication time of document i, t now is the current time, γ is the time decay rate, T is the time scale function; S i is the ratio of the monthly search volume of document i to the maximum search volume of documents in the database; J i is the ratio of the monthly page views of document i to the maximum page views of documents in the database;
[0016] S43: The second convolutional neural network model finally outputs the comprehensive relevance of each related document;
[0017] S5: Display the documents to users according to the comprehensive relevance ranking.
[0018] Preferably, the scientific literature is a paper or patent text; the preprocessing of the uploaded scientific literature or query statement includes word segmentation, removal of stop words and stem extraction.
[0019] Preferably, the S2: extracting semantic features from the preprocessed text to obtain a final feature vector F comprises:
[0020] S21: Use the BERT model to extract the global feature vector F1 from the preprocessed text;
[0021] S22: Input the output of the BERT model into the first convolutional neural network model to obtain a local feature vector F2;
[0022] S23: The global feature vector F1 output by the BERT model is merged with the local feature vector F2 output by the first convolutional neural network model to generate a final feature vector F.
[0023] Preferably, the S21: extracting the preprocessed text using the BERT model to obtain a global feature vector F1 includes:
[0024] S211: Decompose the preprocessed text into words and generate word embedding, position embedding and paragraph embedding for each word;
[0025] S212: Input the generated word embedding, position embedding, and paragraph embedding into multiple Transformer encoding layers of BERT for calculation;
[0026] S213: Calculate the context of each word through the self-attention mechanism;
[0027] S214: The last layer output of the BERT model is the global feature vector F1.
[0028] Preferably, the S22: inputting the output of the BERT model into the first convolutional neural network model to obtain a local feature vector F2, including: the first convolutional neural network model uses a one-dimensional convolutional neural network to process the global feature vector F1, the size of each convolution kernel is k*d, k is the height of the convolution kernel, that is, the number of consecutive words covered by the convolution kernel, d is the dimension of word embedding, and the activation function used in the convolution operation is the ReLU activation function; a pooling operation is performed after the convolution operation, and average pooling is used to reduce the dimension of the feature; the local feature vector F2 is obtained by the convolution and pooling operations.
[0029] Preferably, the step S3: determining the related documents and their monthly search volume and pageview volume from the database according to the final feature vector F comprises:
[0030] S31: Calculate the similarity between the final feature vector F and the feature vectors of scientific literature in the database;
[0031] S32: using cosine similarity to calculate the similarity between the final feature vector and the feature vector of each document;
[0032] S33: Based on the similarity calculation results, scientific and technological documents with similarity greater than a set threshold are screened out as related documents; and the monthly search volume and pageview volume of the related documents are extracted from the database.
[0033] Preferably, S23: fusing the global feature vector F1 output by the BERT model with the local feature vector F2 output by the first convolutional neural network model to generate a final feature vector F, including connecting the global feature vector F1 and the local feature vector F2 end to end to form the final feature vector F.
[0034] The present application also provides a scientific and technological intelligence correlation analysis system, including:
[0035] A query module receives scientific and technological documents or query statements uploaded by users, and preprocesses the uploaded scientific and technological documents or query statements to obtain preprocessed texts;
[0036] The feature vector acquisition module extracts semantic features from the preprocessed text to obtain the final feature vector F;
[0037] The related document determination module determines the related documents and their monthly search volume and pageview volume from the database according to the final feature vector F;
[0038] A comprehensive relevance calculation module, which uses a second convolutional neural network algorithm to calculate the comprehensive relevance between the scientific and technological literature or query statement uploaded by the user and the determined related literature;
[0039] An input pair generation module forms a pair of the final feature vector F and the feature vector of each associated document in the database as an input pair, where the input pair includes the monthly search volume and pageview volume of the associated document;
[0040] The comprehensive relevance score calculation module inputs all input pairs into the second convolutional neural network model. The second convolutional neural network model uses an improved Sigmoid activation function to perform convolution operations, performs maximum pooling operations on the output after convolution, and combines the pooling results of all convolution kernels through a fully connected layer to calculate the comprehensive relevance score.
[0041] The improved Sigmoid activation function is as follows:
[0042]
[0043] Among them, f(x) is the output of the Sigmoid activation function, x is the output of the convolution operation, and w(t i ) is the time decay weight factor, t i is the publication time of document i, t now is the current time, γ is the time decay rate, T is the time scale function; S i is the ratio of the monthly search volume of document i to the maximum search volume of documents in the database; J i is the ratio of the monthly page views of document i to the maximum page views of documents in the database;
[0044] Comprehensive relevance output module: the second convolutional neural network model finally outputs the comprehensive relevance of each related document;
[0045] The related document sorting and display module displays the related documents to the user according to the comprehensive relevance sorting.
[0046] Preferably, the feature vector acquisition module includes:
[0047] Global feature vector F1 acquisition module: Use the BERT model to extract the global feature vector F1 from the preprocessed text;
[0048] Local feature vector F2 acquisition module: input the output of the BERT model into the first convolutional neural network model to obtain the local feature vector F2;
[0049] Final feature vector F acquisition module: The global feature vector F1 output by the BERT model is merged with the local feature vector F2 output by the first convolutional neural network model to generate the final feature vector F.
[0050] Preferably, the global feature vector F1 acquisition module includes:
[0051] Decomposition module: decomposes the preprocessed text into words and generates word embedding, position embedding and paragraph embedding for each word;
[0052] Encoding layer calculation module: input the generated word embedding, position embedding, and paragraph embedding into multiple Transformer encoding layers of BERT for calculation;
[0053] Contextual relationship calculation module: calculates the contextual relationship of each word through the self-attention mechanism;
[0054] BERT model output module: The last layer output of the BERT model is the global feature vector F1.
[0055] The present invention provides a method and system for analyzing the correlation of scientific and technological information, which can achieve the following beneficial technical effects:
[0056] 1. The present invention provides a method and system for analyzing the relevance of scientific and technological information. The method first receives the scientific and technological documents or query statements uploaded by the user, and performs preprocessing to obtain the preprocessed text; secondly, the semantic features of the preprocessed text are extracted to obtain the final feature vector F; then the related documents are determined from the database according to the final feature vector F; the second convolutional neural network algorithm is used again to calculate the comprehensive relevance between the scientific and technological documents or query statements uploaded by the user and the determined related documents; finally, the documents are displayed to the user according to the comprehensive relevance sorting. The present application determines the related documents according to the final feature vector F, and then uses the second convolutional neural network algorithm to calculate the comprehensive relevance of the related documents, and then displays the documents to the user. By adopting the improved Sigmoid activation function to consider the monthly search volume and pageview volume, the accuracy and efficiency of document screening are greatly improved, and the user experience is greatly increased.
[0057] 2. The present invention inputs all input pairs into the second convolutional neural network model, and the second convolutional neural network model uses the improved Sigmoid activation function to perform convolution operation, performs maximum pooling operation on the output after convolution, and combines the pooling results of all convolution kernels through the fully connected layer to calculate the comprehensive correlation score; the improved Sigmoid activation function converts w(t i ) The time decay weight factor, the publication time of the document, the monthly search volume and the page views of the document are taken into consideration, which greatly improves the accuracy of document screening and enhances user satisfaction.
[0058] 3. The present invention uses the BERT model to extract the global feature vector from the preprocessed text, inputs the output of the BERT model into the first convolutional neural network model to obtain the local feature vector, and fuses the global feature vector F1 output by the BERT model with the local feature vector F2 output by the first convolutional neural network model to generate the final feature vector F, which fully considers the local feature vector and the global feature vector, greatly improves the accuracy and efficiency of document screening, and greatly increases the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0060] Figure 1 It is a schematic diagram of the steps of a method for analyzing the correlation of scientific and technological information of the present invention;
[0061] Figure 2 It is a schematic diagram of a scientific and technological information correlation analysis system of the present invention. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] Embodiment 1:
[0064] In view of the above problems mentioned in the prior art, in order to solve the above technical problems, as shown in the attached Figure 1 As shown: This application provides a method for analyzing the correlation of scientific and technological intelligence, including the steps of:
[0065] S1: Receive the scientific and technological literature or query statement uploaded by the user, and preprocess the uploaded scientific and technological literature or query statement to obtain the preprocessed text; Receive the scientific and technological literature or query statement uploaded by the user: In some embodiments, the user uploads a piece of scientific and technological literature through the system interface, or inputs a query statement. The scientific and technological literature is in PDF, Word document, or plain text format, and the query statement is usually a short text description input by the user. Extract the content of the uploaded literature from the file and convert it into a processable text format. Tokenization is to split the text into individual words or sub-words for subsequent processing. Remove stop words, which are common but meaningless words such as "of", "and", "is", etc., to reduce noise. Stemming or lemmatization is to simplify the words to their root forms to unify different forms of words.
[0066] S2: Extract semantic features from the preprocessed text to obtain the final feature vector F; The BERT model is a powerful natural language processing tool that can capture the deep semantic information of the text through the pre-training and fine-tuning processes. In some embodiments, the preprocessed text is decomposed into words, and these words are converted into an input format that BERT can understand. The BERT model needs to convert the input into word embeddings, position embeddings, and segment embeddings. BERT encoding is to pass the processed input into the encoding layer of the BERT model. BERT uses a multi-layer Transformer structure to perform context encoding on each word and generate the semantic representation of each word. Obtain the global feature vector by extracting the representation of the [CLS] token in the BERT output as the global feature vector. This vector represents the global semantic information of the entire text. In one embodiment, a scientific and technological literature titled "Introduction to Quantum Computing" is processed. Input processing: Decompose the above text into words and generate the input format required by BERT (including word embeddings, position embeddings, segment embeddings).
[0067] BERT encoding is to pass the processed input into the BERT model. After multi-layer Transformer encoding, obtain the context representation of each word. Obtain the global feature vector by extracting the vector corresponding to the [CLS] token in the last layer output of the BERT model. This vector represents the global semantics of the entire literature. Use a convolutional neural network model for local feature extraction. The convolutional neural network (CNN) is an effective tool for capturing local patterns and features in the text. By performing convolution and pooling operations on the features output by BERT, the local semantic information of the text can be extracted.
[0068] In some embodiments, the convolution operation uses the global feature vector output by BERT as input, and extracts local features of the text through a one-dimensional convolution operation. The convolution operation applies multiple convolution kernels, each of which can identify different local patterns; the activation function applies the ReLU activation function to the output of the convolution operation to introduce nonlinear features; the pooling operation pools the output after activation (such as maximum pooling or average pooling) to reduce the dimension and extract important local features; the local feature vector is obtained, that is, the result after pooling is the local feature vector, which represents the important local patterns in the text. After extracting the global features, the contextual representation of each word is obtained.
[0069] In one embodiment, local features are extracted from the representations of these words through CNN: the convolution operation is to perform a convolution operation on the representation of each word to extract the local pattern between consecutive words. When the convolution kernel size is 3, the triple word pattern can be identified; the activation function is to apply the ReLU activation function to make the convolution result more diverse; the pooling operation is to perform a pooling operation on the convolution result, select the most important features in each convolution kernel, and form a reduced feature vector; the local feature vector is obtained, that is, the combination of the pooled vectors is the local feature vector, which represents the key local pattern in the document.
[0070] In one embodiment, when the global and local features are fused, a comprehensive feature vector can be obtained by fusing the global feature vector extracted by BERT with the local feature vector extracted by CNN. This vector will simultaneously reflect the global semantics and local patterns of the text. Simple connection is to connect the global feature vector and the local feature vector end to end to form a new high-dimensional vector; weighted connection is to apply different weights to the global and local feature vectors before connection, and then connect them to highlight important features; attention mechanism is to use the attention mechanism to dynamically adjust the weights of global and local features according to the importance of the features, and then combine them. In one embodiment, the global feature vector (dimension is 768) and the local feature vector (dimension is 256) are connected end to end to obtain the final feature vector (dimension is 1024).
[0071] S3: Determine the relevant documents and their monthly search volume and page views from the database according to the final feature vector F;
[0072] S4: Calculate the comprehensive correlation between the scientific and technological documents or query statements uploaded by the user and the determined related documents using the second convolutional neural network algorithm; In some embodiments, extract the document feature vector from the database, first construct or query the feature vector library, that is, store the feature vectors of a large number of scientific and technological documents in the database in advance, and these vectors are extracted and stored by a method similar to step S2, and the real-time calculation is that for newly added documents or dynamic queries, the feature vectors can be calculated in real time when needed. There is a series of processed scientific and technological documents in the database, and each document has been converted into a feature vector. The feature vector of document A is Va, the feature vector of document B is Vb, and the feature vector of document C is Vc; Secondly, calculate the similarity between the feature vector of the uploaded document or query statement and the feature vector of the document in the database, and the similarity measurement method, that is, the commonly used similarity measurement method, includes cosine similarity, Euclidean distance or Manhattan distance. Cosine similarity is often used for similarity comparison of high-dimensional text feature vectors; The feature vector of the document uploaded by the user is Vup, and the similarity between Vup and the feature vectors of documents A, B, and C needs to be calculated. According to the result of the similarity calculation, it can be determined which documents are most relevant to the uploaded document or query statement. Preferably, the associated documents are screened according to the similarity, and a similarity threshold is set, that is, a similarity threshold is set according to specific needs, and documents with similarity higher than the threshold are screened out as associated documents; similarity sorting is to sort these associated documents from high to low according to similarity, so as to give priority to displaying the most relevant documents. Again, the monthly search volume and pageviews of the associated documents are extracted. The monthly search volume is to query the monthly search volume of each associated document from the database, and the search volume is recorded and stored by the system, reflecting the number of searches for the document by the user within a certain period of time. The monthly pageviews is to query the monthly pageviews of each associated document from the database, and the pageviews are also recorded and stored by the system, showing the frequency of users viewing the document within a specific time period. In one embodiment, after selecting document C and document A as associated documents, their monthly search volume and pageviews are extracted from the database, and the monthly search volume of document C is 1200 times and the pageviews are 500 times; the monthly search volume of document A is 900 times and the pageviews are 400 times.
[0073] S41: The final feature vector F and the feature vector of each associated document in the database are paired as an input pair, wherein the input pair includes the monthly search volume and pageview volume of the associated document; in some embodiments, the final feature vector and the associated document feature vector are obtained, and the final feature vector is the final comprehensive feature vector that has been extracted and generated from the scientific and technological documents or query statements uploaded by the user through step S2. The associated document feature vector is the associated documents similar to the uploaded document or query statement that have been determined through step S3, and the feature vectors of these documents have been obtained from the database. Get the monthly search volume and pageviews. Extract the monthly search volume and pageviews of each related document from the database. These data reflect the popularity of the document within a specified time. The monthly search volume and monthly pageviews of document A, document B, and document C form input pairs. The final feature vector and the feature vector of each related document form input pairs. The composition of the input pair is to combine the final feature vector of the document uploaded by the user with the feature vector of each related document to form an input pair. In addition, the input pair also includes the monthly search volume and pageviews of the related document. This information will be used for subsequent convolutional neural network processing and correlation calculation. These input pairs will be input into the second convolutional neural network model, and the model will use the improved Sigmoid activation function to process these pairs. By performing convolution and pooling operations on the input pairs, the comprehensive correlation score of each related document is calculated.
[0074] In some embodiments, the input pairs are as follows: Input pair 1: final feature vector Vup, feature vector Va of document A, monthly search volume SA of document A = 1200, monthly page views VA of document A = 500. Input pair 2: final feature vector Vup, feature vector Vb of document B, monthly search volume SB of document B = 900, monthly page views VB of document B = 400. Input pair 3: final feature vector Vup, feature vector Vc of document C, monthly search volume SC of document C = 1500, monthly page views VC of document C = 600; These input pairs contain the feature vectors of each associated document and their monthly search volume and page views information, which will be input into the convolutional neural network model in subsequent steps to calculate the comprehensive association.
[0075] S42: All input pairs are input into the second convolutional neural network model. The second convolutional neural network model uses an improved Sigmoid activation function to perform a convolution operation, performs a maximum pooling operation on the output after convolution, and combines the pooling results of all convolution kernels through a fully connected layer to calculate a comprehensive correlation score.
[0076] The improved Sigmoid activation function is as follows:
[0077]
[0078] Among them, f(x) is the output of the Sigmoid activation function, x is the output of the convolution operation, and w(t i ) is the time decay weight factor, t i is the publication time of document i, t now is the current time, in days, months or years. The γ time decay rate is a parameter used to control the rate at which the time decay factor changes over time. Its value determines how the timeliness of the document affects its weight. The larger the value, the faster the influence of the document decays over time. According to experience, the weight of the document should be reduced to 50% of the original one year later. T is the time scale function of one year. S i is the ratio of the monthly search volume of document i to the maximum search volume of documents in the database; J i is the ratio of the monthly page views of document i to the maximum page views of documents in the database;
[0079] In some embodiments, multiple input pairs have been generated, which include document feature vectors uploaded by users, feature vectors of associated documents, and monthly search volume and pageviews of associated documents. These data are the basis for entering the second convolutional neural network model, and the user document feature vector is merged with the feature vector of each associated document to form a high-dimensional input vector, which also includes the monthly search volume and pageviews of associated documents as additional information for model calculation. Through the convolution operation, local features and patterns are extracted from the merged input vector, and the convolution kernel is applied to the merged high-dimensional input vector. These convolution kernels scan local areas of the input vector and extract features with similar patterns. The convolution operation helps capture local information of different parts of the input vector, such as semantic relationships and interactions between features. The merged vector (the connection between the user document features and the associated document features) is input into the convolutional neural network. The convolution kernel slides on this vector and calculates the weighted sum of the local area. This process is similar to identifying edges and shapes in image processing, but is used here to identify text features and patterns. Improved Sigmoid activation aims to introduce nonlinear features and dynamically adjust the importance of features by combining factors such as time decay, semantic attention, monthly search volume and page views. The output of the convolution operation is input into an improved Sigmoid activation function. The activation function not only takes into account the basic value of the convolution output, but also combines the time factor, semantic importance and popularity of the associated documents to adjust the output value. The time decay weight factor takes into account the publication time of the document. Newer documents are more relevant, so the time decay weight factor assigns higher weights to these documents. The monthly search volume and page views of the document reflect its popularity and influence. The activation function takes these factors into account and increases the weight of documents with high search volume and page views. Through this improved Sigmoid activation function, the comprehensive importance of the document in terms of time, semantics and popularity can be more accurately reflected. The purpose of the maximum pooling operation is to extract the most important features and reduce the dimension of the data through the pooling operation. The output after the activation function is processed will be input to the pooling layer for the maximum pooling operation. The maximum pooling operation will select the maximum value from the output of each convolution kernel to extract the most important features. In the output of each convolution kernel, the maximum pooling operation will scan all feature values and select the highest value. This process helps to concentrate on retaining the features that contribute most to the overall relevance while reducing the amount of data and computational complexity. The result after pooling is a smaller feature vector that contains the most important information extracted from the convolution operation. The purpose of the fully connected layer calculation is to convert the pooled feature vector into a comprehensive relevance score. The pooled feature vector is input to the fully connected layer. The fully connected layer uses linear transformation and nonlinear activation to finally calculate a relevance score. The fully connected layer multiplies the pooled feature vector with the learned weights and adds a bias.This process is similar to mapping multi-dimensional features to a single score, which represents the comprehensive correlation between the documents or query statements uploaded by the user and specific related documents. The fully connected layer learns the optimal weights and biases through the training process so that the score can accurately reflect the correlation between the documents.
[0080] S43: The second convolutional neural network model finally outputs the comprehensive relevance of each related document;
[0081] In some embodiments, the second convolutional neural network (CNN2) structure purpose; comprehensively consider the input pair (feature vectors of user documents and associated documents) and calculate the comprehensive relevance between documents. Input layer: Receive input pairs consisting of the final feature vector and the feature vector of the associated document, as well as information on monthly search volume and pageviews. The input vector is high-dimensional, 2048 dimensions (merged feature vector) plus search volume and pageview information. Convolution layer: The convolution kernel uses multiple one-dimensional convolution kernels to process the input pair and extract comprehensive features; the convolution operation applies the convolution kernel to the input vector to generate a feature map to capture the complex patterns and relationships in the input pair. Convolution layer parameters: The convolution kernel size 3*2048 means that the convolution kernel covers a local area in the input pair. The number of convolution kernels will have multiple convolution kernels (such as 128 or 256) to capture a variety of comprehensive features. Activation layer: Apply the improved Sigmoid activation function, introduce factors such as time decay, semantic attention, monthly search volume and pageviews. This activation function adjusts the value of the convolution layer output so that important features have a greater impact in subsequent processing. The time decay weight factor takes into account the timeliness of the document, with newer documents having higher weights. Monthly search volume and page views: Adjust the scaling of features so that popular documents have a higher weight in the relevance calculation. Pooling layer: Max pooling selects the maximum value from the activated feature map, reduces the dimension of the feature map, and extracts the most important features. The pooling operation further reduces the data dimension and retains key comprehensive features. Pooling layer parameters: The pooling window size is 2 or 3, indicating the area covered by the pooling operation. Fully connected layer: The pooled feature vector is input into the fully connected layer for linear transformation and nonlinear activation. The fully connected layer calculates the final comprehensive relevance score through the learned weights and biases. Fully connected layer function calculates the comprehensive relevance score: through the learning of weights and biases, a numerical value is generated to represent the comprehensive relevance between the user document and the related documents. The input is a 2048-dimensional feature vector (the combined vector of user documents and related documents). The structure of CNN2 is as follows: input layer: 2048 dimensions; convolution layer: 128 3*2048 convolution kernels -> output 128 feature maps; improved Sigmoid activation layer: apply the improved Sigmoid activation function to the 128 feature maps, considering time and popularity; maximum pooling layer: the pooling window size is 2 -> output 64 reduced feature maps; fully connected layer: flatten the pooled feature map to form a 64*1024 vector -> calculate the final comprehensive relevance score.
[0082] The first convolutional neural network (CNN1) is mainly used to extract local features from the feature vector of a single document and identify local patterns and semantic information in the text. Its structure includes input layer, convolution layer, ReLU activation layer, maximum pooling layer and output layer. The second convolutional neural network (CNN2) is used to process the input pairs consisting of the feature vectors of user documents and related documents, and comprehensively consider factors such as time decay, semantic attention, monthly search volume and pageviews to calculate the comprehensive correlation between documents. Its structure includes input layer, convolution layer, improved Sigmoid activation layer, maximum pooling layer and fully connected layer.
[0083] S5: Display the documents to the user according to the comprehensive relevance ranking. In one embodiment, the user uploaded a document on "quantum computing", and the system found the following related documents from the database. The comprehensive relevance score of each document is as follows: Document A: comprehensive relevance score 0.85, Document B: comprehensive relevance score 0.75, Document C: comprehensive relevance score 0.90, Document D: comprehensive relevance score 0.65. These scores were calculated by CNN2 in the previous section, reflecting the relevance of these documents to the documents uploaded by the user. Sort all related documents from high to low according to the comprehensive relevance score. Use a sorting algorithm (such as quick sort, merge sort, etc.) to sort the documents by score. The sorted result is an ordered list, and the document with the highest score is ranked first, indicating that it has the highest relevance to the document or query statement uploaded by the user. Sort the documents according to the comprehensive relevance score: Document C: 0.90, Document A: 0.85, Document B: 0.75, Document D: 0.65. The sorting results show that Document C is most relevant to the document uploaded by the user, and Document D has the lowest relevance. Generate user-friendly display information for each document. Obtain information such as the title, abstract, author, publication date, and comprehensive relevance score of each document. Organize this information into a structured format for easy display on the user interface. Display details include: the title of the document, which allows users to quickly understand the subject of the document; the abstract of the document, which provides a brief overview to help users evaluate the content and relevance of the document; the author information of the document, which shows the source and researchers of the document; the publication date of the document, which helps users understand the timeliness of the document; and display the relevance score of the document to let users know the relevance of the document to the query.
[0084] In some embodiments, the scientific and technological literature is a paper or a patent text; the preprocessing of the uploaded scientific and technological literature or query statement includes word segmentation, stop word removal, and stemming. Among them, a paper is an academic article published in an academic journal or conference, which structurally reports research results, including parts such as title, abstract, introduction, methods, results, discussion, and references. A patent text is a technical document used to protect invention rights, including parts such as abstract, background art, invention content, specific implementation manners, and claims. A paper is a structured text containing rich domain terms and detailed research descriptions. A patent text is a legal document containing detailed information on technical descriptions and right protection. The preprocessing includes the following key steps: Word segmentation: decomposing the text into individual words or sub-words. Word segmentation is a basic step in text preprocessing, especially important for Chinese and other languages that do not separate words by spaces. Stop word removal: Stop words refer to common words considered to have no practical meaning in text processing, such as "de", "he", "shi", etc. Removing these words can reduce noise, retain meaningful words, and improve the efficiency and accuracy of analysis. Stemming or lemmatization: Stemming simplifies words to their root forms, and lemmatization converts words to their standard forms. This step helps to unify words in different forms, making the analysis process more simplified and consistent.
[0085] In some embodiments, S2: extracting semantic features from the preprocessed text to obtain the final feature vector F, including:
[0086] S21: Use the BERT model to extract the global feature vector F1 from the preprocessed text; the input of the BERT model includes word embedding, position embedding and paragraph embedding, which help the model understand the semantics and structure of the text. First, decompose the preprocessed text into words or subwords that BERT can handle. The text "Quantum computing is developing rapidly" is decomposed into: "quantum", "computing", "being", "rapidly", "developing"; add the [CLS] tag at the beginning of the text to represent the global features of the entire sequence; add the [SEP] tag at the end of the text to separate different sentences (even if there is only one sentence); generate word embeddings Use BERT's vocabulary to convert the decomposed words into word embeddings. Each word embedding is a high-dimensional vector (768 dimensions) that represents the position of the word in the semantic space. Example: The word "quantum" is mapped to a 768-dimensional vector. Generate position embedding: Generate a position embedding for each word to indicate its position in the sequence. This helps the model understand the word order. Example: The position embedding corresponds to the position of each word in the sentence, and 'quantum' is the second position. Generate Paragraph Embeddings: Generate paragraph embeddings for the entire sentence to help distinguish words from different paragraphs. For a single paragraph of text, the paragraph identifier is the same for all words. Example: In a single paragraph, all words have the same paragraph embedding. Input Transformer Encoder Layer of BERT Model Input BERT Model: Input the generated final embeddings into the encoder layer of BERT model. BERT uses multiple Transformer Encoder layers to process these embeddings, each layer contains self-attention mechanism and feed-forward neural network. Self-Attention Mechanism: Self-attention mechanism helps BERT capture the relationship between each word and other words and understand the context. The representation of each word is updated by interacting with all other words in the sequence. Hierarchical Encoding: Through multiple layers of encoding, BERT gradually adjusts and enhances the representation of words in each layer. BERT Basic has 12 layers of encoders, and BERT Large has 24 layers of encoders. Extract Global Feature Vector: The output of BERT model is a vector representation of a sequence, including the context representation of each word and the global representation of the entire sequence. Use the vector of the first position of BERT output (corresponding to the [CLS] tag) as the global feature vector of the entire text. Using global feature vectors for text representation: The vector output by the [CLS] tag of the BERT model is considered to be the global feature vector of the text, which captures the semantic information of the entire input sequence. This vector can be used for various downstream tasks, such as classification, relevance calculation, etc. Subsequent processing: In the literature relevance analysis, this global feature vector will be used as input and further combined with other models (such as convolutional neural networks) to extract more information or calculate relevance.
[0087] S22: Input the output of the BERT model to the first convolutional neural network model to obtain the local feature vector F2; in one implementation, the first convolutional neural network (CNN1) structure extracts local features from the preprocessed text features to capture local patterns and semantic information in the text. Input layer: Receives the global feature vector output by the BERT model; the input vector is a high-dimensional, 1024-dimensional word vector. Convolution layer: The convolution kernel uses multiple one-dimensional convolution kernels (such as lengths of 3, 5, 7, etc.) to process the input vector in a sliding window manner to extract local features. In the convolution operation, each convolution kernel scans a portion of the input vector and calculates the weighted sum to generate a feature map. The output generates multiple feature maps, each of which represents the response of the input vector under different convolution kernels. The convolution kernel size of 3*1024 means that the convolution kernel covers the local features of 3 consecutive words. There will be multiple convolution kernels (such as 64 or 128) to capture multiple local patterns. The ReLU activation function is applied to increase nonlinearity and improve the expressiveness of the model. The ReLU activation function returns negative values in the output of the convolution layer to zero, retains positive values, and increases sparsity. Max pooling selects the maximum value from the feature map after convolution, reduces the dimension of the feature map, and retains important features. The pooling operation reduces the data dimension, alleviates overfitting, and extracts important local features. The pooling window size is 2 or 3, indicating the area covered by the pooling operation. The output layer flattens the pooled feature map to form a local feature vector; this local feature vector represents the key local pattern extracted from the input text. In one embodiment, the first convolutional neural network structure is that the input is a 1024-dimensional feature vector, and the structure is as follows: input layer: 1024 dimensions; convolution layer: 64 3*1024 convolution kernels -> output 64 feature maps; ReLU activation layer: apply ReLU to 64 feature maps; max pooling layer: pooling window size is 2 -> output 32 reduced feature maps; output layer: flatten the feature map to form a 32*512 local feature vector.
[0088] S23: The global feature vector F1 output by the BERT model is merged with the local feature vector F2 output by the first convolutional neural network model to generate a final feature vector F.
[0089] In some embodiments, the step S21: extracting the preprocessed text using the BERT model to obtain a global feature vector F1 includes:
[0090] S211: Decompose the preprocessed text into words and generate word embedding, position embedding and paragraph embedding for each word;
[0091] S212: Input the generated word embedding, position embedding, and paragraph embedding into multiple Transformer encoding layers of BERT for calculation;
[0092] S213: Calculate the contextual relationship of each word through the self-attention mechanism; the self-attention mechanism includes the following main steps: generate query, key and value vectors, calculate the attention score, apply the Softmax function, and perform weighted summation; Step 1: Generate query, key and value vectors, and each input word is converted into three vectors: query, key and value. These vectors are obtained by linearly transforming the input embedding. Step 2: Calculate the attention score, and use the dot product between the query vector and the key vector to measure the relevance between the word and other words. The result of the dot product is called the attention score; Step 3: Apply the Softmax function, and convert the attention score into attention weights through the Softmax function. These weights are positive numbers and the sum is 1. Step 4: Weighted sum value vectors, and use the attention weights to perform weighted summation on the value vectors to obtain the contextual representation of each word. The contextual representation is the weighted average of the value vectors, reflecting the importance and relevance of the word in the sequence. After this processing, these weights can be used to perform weighted summation on the value vectors. The main function of the self-attention mechanism is to capture the contextual relationship. The self-attention mechanism helps the model capture the relationship between each word and other words, no matter how far away they are. This mechanism allows the BERT model to understand complex dependencies and semantic structures. Formation of global feature vector: Through the self-attention mechanism, BERT can effectively combine the information of the entire sequence to form a global feature representation of the input sequence. Ultimately, the vector of the [CLS] tag position output by BERT is used as the global feature vector of the entire sequence.
[0093] S214: The last layer output of the BERT model is the global feature vector F1.
[0094] In some embodiments, the S22: inputting the output of the BERT model into the first convolutional neural network model to obtain a local feature vector F2, including: the first convolutional neural network model uses a one-dimensional convolutional neural network to process the global feature vector F1, the size of each convolution kernel is k*d, k is the height of the convolution kernel, that is, the number of consecutive words covered by the convolution kernel, d is the dimension of word embedding, and the activation function used in the convolution operation is the ReLU activation function; a pooling operation is performed after the convolution operation, and average pooling is used to reduce the dimension of the feature; the local feature vector F2 is obtained through convolution and pooling operations.
[0095] In some embodiments, the step S3: determining related documents from the database according to the final feature vector F, and obtaining the monthly search volume and pageview volume of the related documents, includes:
[0096] S31: Calculate the similarity between the final feature vector F and the feature vectors of scientific literature in the database;
[0097] S32: using cosine similarity to calculate the similarity between the final feature vector and the feature vector of each document;
[0098] S33: Based on the similarity calculation results, scientific and technological documents with similarity greater than a set threshold are screened out as related documents; and the monthly search volume and pageview volume of the related documents are extracted from the database.
[0099] In some embodiments, S23: fusing the global feature vector F1 output by the BERT model with the local feature vector F2 output by the first convolutional neural network model to generate a final feature vector F, including connecting the global feature vector F1 and the local feature vector F2 end to end to form the final feature vector F.
[0100] This application also provides a scientific and technological intelligence correlation analysis system, such as Figure 2 As shown, the system hardware consists of servers / computer clusters, data storage devices, network connection devices, and user interface devices. Among them, the server or computer cluster is the core computing unit of the system, responsible for processing user requests and performing complex computing tasks such as preprocessing, feature extraction, correlation calculation, etc., including components: Central Processing Unit (CPU): Processes most of the computing tasks of the system. Graphics Processing Unit (GPU): Used for training and reasoning of deep learning models, accelerating processing such as BERT models and convolutional neural network calculations. Memory (RAM): Used to temporarily store runtime data and calculate intermediate results. Storage: Stores operating systems, applications, and temporary data. Servers / computer clusters are interconnected through an internal network (high-speed Ethernet) to support distributed computing and data processing.
[0101] Data storage devices are used to store large amounts of scientific and technological literature data, user data, feature vector libraries, intermediate calculation results, and model parameters. Components include: Database servers: used to store structured data, such as literature metadata, user query history, etc. File storage systems: used to store unstructured data, such as the full text content of literature, PDF files, etc. Distributed storage systems: used to store large-scale data and model files, such as HDFS or AWS S3. Data storage devices are connected to servers / computer clusters via high-speed networks for fast access and processing of data.
[0102] Network connection devices ensure efficient communication between components within the system and between the system and external networks. Components include: Routers: manage network traffic and ensure data transmission within and outside the system. Switches: connect servers, storage devices and other hardware to provide high-speed data transmission within the local area network (LAN). Firewalls: protect the system from external threats and ensure the security of data and computing. Server / computer clusters and data storage devices are connected through switches to form a local area network. Routers connect the local area network to the external Internet, and firewalls protect the entire network architecture.
[0103] The user interface device is used for interaction between the user and the system, supporting users to upload documents, query and view analysis results. The components include: User terminal device: including the user's personal computer, laptop, tablet or smart phone. Browser or client software: users access the system interface through these software and perform interactive operations. Display device: used to display the analysis results of the system, such as a monitor or mobile phone screen. Input device: users interact with the system through a keyboard, mouse or touch screen. The user terminal device is connected to the system's server through the Internet, and users can access the system from anywhere through a browser or dedicated client software.
[0104] In one embodiment, the user uploads scientific literature or inputs a query statement through a user interface device (such as a browser or client software). The user interface device is connected to the server of the system through the Internet and sends a user request. The server / computer cluster receives the user request and calls the corresponding processing module to perform document preprocessing, feature extraction and correlation calculation. The calculation task involves calling the BERT model and convolutional neural network model, which are executed on a server with GPU acceleration. The server accesses the data storage device to retrieve the required document data and feature vectors. The calculation results and intermediate data can be temporarily stored in the storage device for subsequent use or display. After the calculation is completed, the system returns the results to the user through the user interface device. The user can view the results of the document correlation analysis, and the system can display relevant document lists, correlation scores and other information.
[0105] A query module receives scientific and technological documents or query statements uploaded by users, and preprocesses the uploaded scientific and technological documents or query statements to obtain preprocessed texts;
[0106] The feature vector acquisition module extracts semantic features from the preprocessed text to obtain the final feature vector F;
[0107] The related document determination module determines the related documents and their monthly search volume and pageview volume from the database according to the final feature vector F;
[0108] A comprehensive relevance calculation module, which uses a second convolutional neural network algorithm to calculate the comprehensive relevance between the scientific and technological literature or query statement uploaded by the user and the determined related literature;
[0109] An input pair generation module forms a pair of the final feature vector F and the feature vector of each associated document in the database as an input pair, where the input pair includes the monthly search volume and pageview volume of the associated document;
[0110] The comprehensive relevance score calculation module inputs all input pairs into the second convolutional neural network model. The second convolutional neural network model uses an improved Sigmoid activation function to perform convolution operations, performs maximum pooling operations on the output after convolution, and combines the pooling results of all convolution kernels through a fully connected layer to calculate the comprehensive relevance score.
[0111] The improved Sigmoid activation function is as follows:
[0112]
[0113] Among them, f(x) is the output of the Sigmoid activation function, x is the output of the convolution operation, and w(t i ) is the time decay weight factor, t i is the publication time of document i, t now is the current time, γ is the time decay rate, T is the time scale function; S i is the ratio of the monthly search volume of document i to the maximum search volume of documents in the database; J i is the ratio of the monthly page views of document i to the maximum page views of documents in the database;
[0114] Comprehensive relevance output module: the second convolutional neural network model finally outputs the comprehensive relevance of each related document;
[0115] The related document sorting and display module displays the related documents to the user according to the comprehensive relevance sorting.
[0116] In some embodiments, the feature vector acquisition module includes:
[0117] Global feature vector F1 acquisition module: Use the BERT model to extract the global feature vector F1 from the preprocessed text;
[0118] Local feature vector F2 acquisition module: input the output of the BERT model into the first convolutional neural network model to obtain the local feature vector F2;
[0119] Final feature vector F acquisition module: The global feature vector F1 output by the BERT model is merged with the local feature vector F2 output by the first convolutional neural network model to generate the final feature vector F.
[0120] In some embodiments, the global feature vector F1 acquisition module includes:
[0121] Decomposition module: decomposes the preprocessed text into words and generates word embedding, position embedding and paragraph embedding for each word;
[0122] Encoding layer calculation module: input the generated word embedding, position embedding, and paragraph embedding into multiple Transformer encoding layers of BERT for calculation;
[0123] Contextual relationship calculation module: calculates the contextual relationship of each word through the self-attention mechanism;
[0124] BERT model output module: The last layer output of the BERT model is the global feature vector F1.
[0125] The present invention provides a method and system for analyzing the correlation of scientific and technological information, which can achieve the following beneficial technical effects:
[0126] 1. The present invention provides a method and system for analyzing the relevance of scientific and technological information. The method first receives the scientific and technological documents or query statements uploaded by the user, and performs preprocessing to obtain the preprocessed text; secondly, the semantic features of the preprocessed text are extracted to obtain the final feature vector F; then the related documents are determined from the database according to the final feature vector F; the second convolutional neural network algorithm is used again to calculate the comprehensive relevance between the scientific and technological documents or query statements uploaded by the user and the determined related documents; finally, the documents are displayed to the user according to the comprehensive relevance sorting. The present application determines the related documents according to the final feature vector F, and then uses the second convolutional neural network algorithm to calculate the comprehensive relevance of the related documents, and then displays the documents to the user. By adopting the improved Sigmoid activation function to consider the monthly search volume and pageview volume, the accuracy and efficiency of document screening are greatly improved, and the user experience is greatly increased.
[0127] 2. The present invention inputs all input pairs into the second convolutional neural network model, and the second convolutional neural network model uses the improved Sigmoid activation function to perform convolution operation, performs maximum pooling operation on the output after convolution, and combines the pooling results of all convolution kernels through the fully connected layer to calculate the comprehensive correlation score; the improved Sigmoid activation function converts w(t i ) The time decay weight factor, the publication time of the document, the monthly search volume and the page views of the document are taken into consideration, which greatly improves the accuracy of document screening and enhances user satisfaction.
[0128] 3. The present invention uses the BERT model to extract the global feature vector from the preprocessed text, inputs the output of the BERT model into the first convolutional neural network model to obtain the local feature vector, and fuses the global feature vector F1 output by the BERT model with the local feature vector F2 output by the first convolutional neural network model to generate the final feature vector F, which fully considers the local feature vector and the global feature vector, greatly improves the accuracy and efficiency of document screening, and greatly increases the user experience.
[0129] The above is a detailed introduction to a method and system for analyzing the correlation between scientific and technological intelligence. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. At the same time, for those skilled in the art, according to the ideas and methods of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A method for analyzing the correlation of scientific and technological information, characterized in that: Includes steps: S1: receiving scientific documents or query statements uploaded by users, and preprocessing the uploaded scientific documents or query statements to obtain preprocessed texts; S2: Extract semantic features from the preprocessed text to obtain the final feature vector F; S3: Determine the related documents from the database according to the final feature vector F, and obtain the monthly search volume and pageview volume of the related documents; S4: using the second convolutional neural network algorithm to calculate the comprehensive correlation between the scientific literature or query statement uploaded by the user and the determined related literature; S41: Pair the final feature vector F with the feature vector of each associated document in the database as an input pair, wherein the input pair includes the monthly search volume and pageview volume of the associated document; S42: All input pairs are input into the second convolutional neural network model. The second convolutional neural network model uses an improved Sigmoid activation function to perform a convolution operation, performs a maximum pooling operation on the output after convolution, and combines the pooling results of all convolution kernels through a fully connected layer to calculate a comprehensive correlation score. The improved Sigmoid activation function is as follows: Among them, f(x) is the output of the Sigmoid activation function, x is the output of the convolution operation, and w(t i ) is the time decay weight factor, t i is the publication time of document i, t now is the current time, γ is the time decay rate, T is the time scale function; S i is the ratio of the monthly search volume of document i to the maximum search volume of documents in the database; J i is the ratio of the monthly page views of document i to the maximum page views of documents in the database; S43: The second convolutional neural network model finally outputs the comprehensive relevance of each related document; S5: Display the documents to users according to the comprehensive relevance ranking.
2. A method for analyzing the correlation of scientific and technological information according to claim 1, characterized in that: The scientific and technological documents are papers or patent texts; the preprocessing of the uploaded scientific and technological documents or query statements includes word segmentation, removal of stop words and stem extraction.
3. A method for analyzing the correlation of scientific and technological information according to claim 1, characterized in that: S2: extracting semantic features from the preprocessed text to obtain a final feature vector F, including: S21: Use the BERT model to extract the global feature vector F1 from the preprocessed text; S22: Input the output of the BERT model into the first convolutional neural network model to obtain a local feature vector F2; S23: The global feature vector F1 output by the BERT model is merged with the local feature vector F2 output by the first convolutional neural network model to generate a final feature vector F.
4. A method for analyzing the correlation of scientific and technological information as claimed in claim 3, characterized in that: S21: extracting the preprocessed text using the BERT model to obtain a global feature vector F1, including: S211: Decompose the preprocessed text into words and generate word embedding, position embedding and paragraph embedding for each word; S212: Input the generated word embedding, position embedding, and paragraph embedding into multiple Transformer encoding layers of BERT for calculation; S213: Calculate the context of each word through the self-attention mechanism; S214: The last layer output of the BERT model is the global feature vector F1.
5. A method for analyzing the correlation of scientific and technological information as claimed in claim 3, characterized in that: The S22: inputting the output of the BERT model into the first convolutional neural network model to obtain a local feature vector F2, including: the first convolutional neural network model uses a one-dimensional convolutional neural network to process the global feature vector F1, the size of each convolution kernel is k*d, k is the height of the convolution kernel, that is, the number of consecutive words covered by the convolution kernel, d is the dimension of word embedding, and the activation function used in the convolution operation is the ReLU activation function; a pooling operation is performed after the convolution operation, and average pooling is used to reduce the dimension of the feature; the local feature vector F2 is obtained through the convolution and pooling operations.
6. A method for analyzing the correlation of scientific and technological information according to claim 1, characterized in that: S3: determining related documents from the database according to the final feature vector F, including: S31: Calculate the similarity between the final feature vector F and the feature vectors of scientific literature in the database; S32: using cosine similarity to calculate the similarity between the final feature vector and the feature vector of each document; S33: Based on the similarity calculation results, scientific and technological documents with similarity greater than a set threshold are screened out as related documents; and the monthly search volume and pageview volume of the related documents are extracted from the database.
7. A method for analyzing the correlation of scientific and technological information as claimed in claim 3, characterized in that: The S23: fusing the global feature vector F1 output by the BERT model with the local feature vector F2 output by the first convolutional neural network model to generate a final feature vector F, including connecting the global feature vector F1 and the local feature vector F2 end to end to form the final feature vector F.
8. A scientific and technological information correlation analysis system, characterized in that: include: A query module receives scientific and technological documents or query statements uploaded by users, and preprocesses the uploaded scientific and technological documents or query statements to obtain preprocessed texts; The feature vector acquisition module extracts semantic features from the preprocessed text to obtain the final feature vector F; The related document determination module determines the related documents from the database according to the final feature vector F, and obtains the monthly search volume and pageview volume of the related documents; A comprehensive relevance calculation module, which uses a second convolutional neural network algorithm to calculate the comprehensive relevance between the scientific and technological literature or query statement uploaded by the user and the determined related literature; An input pair generation module forms a pair of the final feature vector F and the feature vector of each associated document in the database as an input pair, where the input pair includes the monthly search volume and pageview volume of the associated document; The comprehensive relevance score calculation module inputs all input pairs into the second convolutional neural network model. The second convolutional neural network model uses an improved Sigmoid activation function to perform convolution operations, performs maximum pooling operations on the output after convolution, and combines the pooling results of all convolution kernels through a fully connected layer to calculate the comprehensive relevance score. The improved Sigmoid activation function is as follows: Among them, f(x) is the output of the Sigmoid activation function, x is the output of the convolution operation, and w(t i ) is the time decay weight factor, t i is the publication time of document i, t now is the current time, γ is the time decay rate, T is the time scale function; S i is the ratio of the monthly search volume of document i to the maximum search volume of documents in the database; J i is the ratio of the monthly page views of document i to the maximum page views of documents in the database; Comprehensive relevance output module: the second convolutional neural network model finally outputs the comprehensive relevance of each related document; The related document sorting and display module displays the related documents to the user according to the comprehensive relevance sorting.
9. A scientific and technological information correlation analysis system as claimed in claim 8, characterized in that: The feature vector acquisition module comprises: Global feature vector F1 acquisition module: Use the BERT model to extract the global feature vector F1 from the preprocessed text; Local feature vector F2 acquisition module: input the output of the BERT model into the first convolutional neural network model to obtain the local feature vector F2; Final feature vector F acquisition module: The global feature vector F1 output by the BERT model is merged with the local feature vector F2 output by the first convolutional neural network model to generate the final feature vector F.
10. A scientific and technological information correlation analysis system as claimed in claim 9, characterized in that: The global feature vector F1 acquisition module includes: Decomposition module: decomposes the preprocessed text into words and generates word embedding, position embedding and paragraph embedding for each word; Encoding layer calculation module: input the generated word embedding, position embedding, and paragraph embedding into multiple Transformer encoding layers of BERT for calculation; Contextual relationship calculation module: calculates the contextual relationship of each word through the self-attention mechanism; BERT model output module: The last layer output of the BERT model is the global feature vector F1.
Citation Information
Patent Citations
User interest graph sequence dynamic management method based on search keywords
CN111488493A
Scientific and technical literature quotation recommendation method based on deep learning
CN113239181A