Multi-dimension-based frontier technology recognition model training method, recognition method and device
Through multi-dimensional cutting-edge technology identification model, combined with statistical meta information and text data, multi-layer perceptron and large language models are used to fusion of features, solving the problem of insufficient understanding of literature content in the existing technology, and achieving efficient and accurate identification of cutting-edge research.
Patent Information
- Application Number
- CN202510771223.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology lacks an in-depth understanding of the literature content in cutting-edge research and identification, and the analysis dimension is single, making it difficult to accurately identify interdisciplinary innovations and emerging technologies, and fail to capture the potential research direction in a timely manner.
The multi-dimensional cutting-edge technology recognition model is adopted, and by acquiring and preprocessing the statistical meta information and text data of scientific and technological literature, combining multi-layer perceptrons and pre-trained large language models, the potential cross attention module is used to fusion of features to generate high-dimensional feature vectors to identify cutting-edge technology.
It has achieved comprehensive and accurate identification of scientific and technological literature, improved the accuracy and timeliness of identification of cutting-edge technology, and can timely capture innovative research directions.
Smart Images

Figure CN120278149A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of literature data processing, and in particular, to a method and device for training a cutting-edge technology recognition model based on multiple dimensions, as well as a recognition method. Background Art
[0002] With the rapid development of technology, the number of scientific research literatures has increased explosively. Identifying cutting-edge research directions is of great significance to researchers, research institutions, and policymakers, etc. It helps to allocate scientific research resources reasonably, determine research priorities, and grasp the trend of scientific and technological development.
[0003] Early cutting-edge identification methods mostly relied on expert experience and qualitative analysis, such as the Delphi method, etc. However, such methods have problems such as strong subjectivity, low efficiency, and difficulty in processing large-scale data. With the continuous progress of computing technology, analysis methods based on quantitative data have gradually emerged and developed continuously.
[0004] Traditional citation analysis method is a commonly used cutting-edge identification means. It evaluates research influence and identifies important achievements by analyzing the citation network between papers, such as the number of citations, co-citation relationships, etc. This method extracts the meta-information of scientific and technological literatures, and uses machine learning methods to predict the number of citations according to statistical features for cutting-edge identification. However, this method has obvious limitations. It overly relies on the citation network structure, lacks in-depth understanding of the literature content, and is difficult to insight into the deep knowledge and innovation points contained in the literature. And there is a lag in the identification of emerging technologies. It needs to wait until the literature accumulates enough citations to be effectively identified, and it is unable to timely capture the emerging and potential research directions.
[0005] The shallow semantic analysis method based on keywords / topics has also been widely used, such as using methods like keyword co-occurrence, topic modeling (such as the LDA algorithm) to identify research hotspots and topic distributions. However, such methods also have many drawbacks. Their semantic understanding only stays at the lexical level, unable to capture deep concept associations, and it is difficult to identify potentially innovative research directions. The analysis dimension is relatively single, and it fails to fully combine other important statistical features, such as time evolution, author network, etc., resulting in an incomplete and inaccurate judgment of cutting-edge research. In addition, it is not sensitive enough to cross-disciplinary innovation research directions. Due to its principle based on keyword matching, when facing concept migration and fusion innovation between different disciplinary fields, it is difficult to effectively identify, thus unable to meet the increasing demand for interdisciplinary intersections in today's scientific research field, and it is limited in effectively integrating multi-source heterogeneous literature features and identifying truly innovative research.
[0006] In view of the problems existing in the prior art, there is an urgent need for a new method that can deeply integrate multi-dimensional information, deeply understand the semantics of literatures, and timely and accurately identify cutting-edge research directions. Summary of the Invention
[0007] In view of this, embodiments of the present invention provide a multi-dimensional training method, recognition method, and device for cutting-edge technology recognition models to eliminate or improve one or more defects existing in the prior art, and solve the problem that the prior art lacks in-depth understanding of document content and has a single analysis dimension during the recognition of scientific and technological frontiers, resulting in insufficient recognition ability.
[0008] One aspect of the present invention provides a multi-dimensional training method for cutting-edge technology recognition models, the method comprising the following steps: Obtain multiple sample data, each sample data comprising statistical meta-information and text data information of a scientific and technological literature, the statistical meta-information comprising publication time, citation quantity, times cited, number of authors, productivity of the first author, and times cited, and the text data information comprising a title and an abstract; Perform preprocessing on each of the sample data, the preprocessing including data screening, data cleaning and standardization of the statistical meta-information, and adding a label indicating whether the scientific and technological literature belongs to cutting-edge technology in combination with the frontier index mark of each sample data to form a first training sample set; the frontier index quantifies the frontier nature of technical literature by combining citation intensity and the contribution level of team members; Obtain an initial neural network model, including a statistical meta-information encoder, a title and abstract encoder, a potential cross-attention module, and a fully connected layer, the statistical meta-information encoder converting the preprocessed statistical meta-information into a high-dimensional meta-information vector based on a multi-layer perceptron, the title and abstract encoder extracting semantic embedding vectors from the vectorized text data information based on a preset embedded large language model, the potential cross-attention module taking the high-dimensional meta-information vector and the semantic embedding vector as inputs and outputting a literature feature vector, and inputting the literature feature vector into the fully connected layer to output a prediction result indicating whether it belongs to cutting-edge technology; training the initial neural network model using the first training sample set, and constructing a loss to update the model parameters based on the deviation between the prediction result and the label to obtain a cutting-edge technology recognition model.
[0009] In some embodiments, performing preprocessing on each of the sample data includes: Screen the sample data according to a preset plurality of sources of scientific and technological literature; Fill in the missing fields of the statistical meta-information in the screened sample data with mean values or mode values and add missing markers for auxiliary recognition; standardize the statistical meta-information using Min-Max scaling or Z-score, detect outliers through Z-score, mark the values exceeding the threshold range as outliers and replace them with boundary values; Convert the publication time into a relative timestamp that is the offset from the current time.
[0010] In some embodiments, adding a label indicating whether the scientific and technological literature belongs to a cutting-edge technology in combination with the leading index marker of each sample data includes: Calculate the leading index of the scientific and technological literature, and the calculation formula is: ; where f represents the leading index, p represents the number of citations, t represents the publication time, represents the contribution ratio of the i-th author of the scientific and technological literature, and N represents the number of authors of the scientific and technological literature; Set leading index thresholds for multiple technical fields respectively, and mark the scientific and technological literatures with the leading index higher than the corresponding leading index threshold as cutting-edge technologies.
[0011] In some embodiments, the embedded large language model is pre-trained using a Transformer natural language processing model, and the pre-training steps include: Collect sample scientific and technological text data information for multiple technical fields, remove noise, invalid content, and sensitive words, perform word segmentation, stop word removal, and stemming, and construct a second training sample set; Based on the second training sample set, the embedded large language model performs masked language modeling or next sentence prediction tasks for pre-training to update the parameters of the embedded large language model.
[0012] In some embodiments, the embedded large language model uses BERT, RoBERTa, GPT, Deepseek, NV_Embed_V2, and bge-en-icl models.
[0013] In some embodiments, the potential cross-attention module performs embedding processing on the semantic embedding vector to obtain a key-value vector, performs embedding processing on the meta-information high-dimensional vector to obtain a query vector, and after fusing the key-value vector and the query vector based on the cross-attention mechanism, obtains the literature feature vector after processing by a feed-forward neural network and a normalization layer; After the title abstract encoder extracts the semantic embedding vector from the vectorized text data information based on a preset embedded large language model, it further includes performing dimensionality reduction mapping on the semantic embedding vector.
[0014] On the other hand, the present invention also provides a multi-dimensional-based cutting-edge technology identification method, and the method includes: Obtain the statistical meta-information and text data information of the scientific and technological literature to be analyzed. The statistical meta-information includes publication time, citation quantity, times cited, number of authors, productivity of the first author, and times cited. The text data information includes title and abstract; Perform standardized processing on the statistical meta-information; Input the standardized statistical meta-information and the text data information into the frontier technology recognition model in the above-mentioned multi-dimensional frontier technology recognition model training method to output the recognition result of whether the scientific and technological literature to be analyzed belongs to the frontier technology.
[0015] On the other hand, the present invention also provides a multi-dimensional frontier technology recognition device, including a processor, a memory, and a computer program / instructions stored on the memory. The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the device implements the steps of the above method.
[0016] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0017] On the other hand, the present invention also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0018] The beneficial effects of the present invention are at least: The multi-dimensional frontier technology recognition model training method, recognition method and device of the present invention introduce multi-dimensional statistical meta-information and text data information to construct a training sample set, and fuse the citation intensity and the contribution level of team members to quantify the frontier nature of technical literature to accurately add labels. The model introduces a multi-layer perceptron to extract high-dimensional features of statistical meta-information, introduces a pre-trained embedded large language model to extract semantic features of the titles and abstracts of scientific and technological documents, and performs feature fusion through a potential cross self-attention module for frontier technology recognition, realizing the deep integration of multi-modal information, enabling the model to comprehensively and accurately grasp the characteristics of scientific and technological literature, and significantly improving the accuracy and timeliness of frontier technology recognition in the application process.
[0019] The additional advantages, objectives, and features of the present invention will be partially described below, and will become partially apparent to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the specification and the drawings.
[0020] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and the above and other objectives achievable with the present invention will be more clearly understood from the following detailed description. Description of the Drawings
[0021] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not constitute a limitation to the present invention. In the drawings: Figure 1 It is a schematic flowchart of a method for training a multi-dimensional cutting-edge technology identification model according to an embodiment of the present invention.
[0022] Figure 2 It is a schematic logical diagram of a method for training a multi-dimensional cutting-edge technology identification model according to another embodiment of the present invention. Detailed Embodiments
[0023] To make the objectives, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not constitute a limitation to the present invention.
[0024] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0025] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0026] Large Language Model (LLM) is a type of deep learning model developed based on the Transformer architecture. Since Vaswani et al. proposed the Transformer architecture in 2017, significant progress has been made in the field of natural language processing along this technical route. Typical LLM models adopt a multi-layer Transformer structure and achieve parallel processing of text sequences through the self-attention mechanism, significantly improving the ability to model long-distance dependencies. Such models usually adopt a two-stage training paradigm: first, self-supervised pre-training is carried out on a large amount of unlabeled text data to learn general language representations; then, fine-tuning or prompt engineering for specific tasks is used to adapt to downstream applications. In recent years, with the continuous expansion of model scale and the continuous increase of training data, LLM has shown powerful language understanding and generation capabilities and has been widely applied in many natural language processing task fields such as machine translation, text summarization, and question answering systems.
[0027] Multidimensional analysis technology originated in the fields of statistics and informetrics and is a method that comprehensively uses a variety of quantitative indicators to evaluate the research object. In the field of scientific research, traditional bibliometric indicators such as impact factor and h-index mainly focus on a single citation dimension. With the development of complex system theory, researchers have begun to incorporate dimensions such as time dimension, network dimension, and semantic dimension into the analysis framework. Time dimension analysis focuses on the evolution trend and life cycle of research topics; network dimension analysis constructs citation networks or collaboration networks and uses graph theory methods to identify key nodes and community structures; semantic dimension analysis uses natural language processing technology to mine the topic distribution and concept associations in texts. These multidimensional quantitative indicators provide a more comprehensive perspective for scientific research evaluation.
[0028] The frontier identification technology aims to discover emerging directions and breakthrough progress in scientific research. In order to overcome the problems of the existing technology lacking in-depth understanding of document content and single analysis dimension, the present invention introduces large language models and multidimensional analysis methods to identify frontier technologies.
[0029] Specifically, the present invention provides a method for training a multidimensional-based frontier technology identification model. As Figure 1 shown, the method includes the following steps S101 to S103: Step S101: Obtain multiple sample data. Each sample data includes statistical meta-information and text data information of scientific and technological literature. The statistical meta-information includes publication time, citation quantity, cited frequency, number of authors, productivity of the first author, and cited frequency. The text data information includes title and abstract.
[0030] Step S102: Perform preprocessing on each sample data. The preprocessing includes data screening, data cleaning and standardization of the statistical meta-information, and adding a label indicating whether the scientific and technological literature belongs to frontier technology in combination with the frontier index mark of each sample data to form a first training sample set; the frontier index quantifies the frontier nature of technical literature by combining citation intensity and the contribution level of team members.
[0031] Step S103: Refer to Figure 2, obtain an initial neural network model, including a statistical meta-information encoder, a title summary encoder, a potential cross-attention module and a fully connected layer. The statistical meta-information encoder converts the preprocessed statistical meta-information into a meta-information high-dimensional vector based on a multi-layer perceptron. The title summary encoder extracts a semantic embedding vector from the vectorized text data information based on a preset embedded large language model. The potential cross-attention module takes the meta-information high-dimensional vector and the semantic embedding vector as input and outputs a document feature vector. The document feature vector is input into the fully connected layer to output a prediction result of whether it belongs to a cutting-edge technology; the first training sample set is used to train the initial neural network model, and the model parameters are updated by constructing a loss based on the deviation between the prediction result and the label to obtain a cutting-edge technology recognition model.
[0032] In step S101, massive scientific and technological literature resources are obtained in batches from large scientific research literature databases such as Semantic Scholar and other platforms to ensure the breadth and diversity of the data, covering literature from different disciplines, fields, and time periods, laying a solid foundation for model training.
[0033] The sample data contains two parts of core information. The first is statistical meta-information, which specifically covers the publication time of the document, accurate to the year, month and day, reflecting the timeliness of the research; the number of citations, that is, the number of times the document cites other documents, reflecting the comprehensiveness of its research and the degree of reference to the achievements of predecessors; the frequency of citations, which measures the degree of attention and recognition of the document by other documents, is a key indicator of academic influence; the number of authors, which indirectly reflects the size and degree of collaboration of the research team; the productivity of the first author, which measures the activity and contribution of the first author in the field by counting the number of scientific research outputs of the first author in a certain period of time; the frequency of citations of the first author, reflects the influence and academic status of the first author's past research results. The second is text data information, including the title and abstract of the document. The title highly summarizes the research topic of the document, and the abstract briefly summarizes the core content, methods, conclusions and other key information of the research, providing key materials for text semantic analysis.
[0034] In step S102, the collected sample data is strictly screened to eliminate low-quality data that does not meet the requirements, such as samples with serious content missing and irrelevant to the subject of scientific literature, and retain high-quality and highly relevant data to ensure the accuracy and reliability of subsequent analysis.
[0035] For each data in the statistical metadata, data cleaning is first performed to remove obvious erroneous values and outliers, such as unreasonable citation numbers and incorrect publication time formats. Then, standardization operations are implemented, and batch normalization and other methods are used to unify the scale of numerical data such as publication time, citation times, and number of authors, eliminating the impact of differences in dimensions and numerical ranges between different feature dimensions on subsequent calculations, thereby improving the stability and efficiency of model training.
[0036] In some embodiments, preprocessing is performed on each of the sample data, including steps S1021 to S1023: Step S1021: Screen the sample data according to a plurality of preset scientific and technological literature sources.
[0037] Step S1022: For the missing fields of the statistical meta-information in the screened sample data, mean filling or mode filling is used for supplementation and a missing marker is added to assist in identification; Min-Max scaling or Z-score is used to standardize the statistical meta-information, and outliers are detected through Z-score. The values exceeding the threshold range are marked as the outliers and then replaced with boundary values.
[0038] Step S1023: Convert the publication time into a relative timestamp of the offset from the current time.
[0039] In step S1021, the sample data is screened according to a plurality of preset scientific and technological literature sources, which include but are not limited to well-known authoritative academic databases in the field (such as Semantic Scholar, Web of Science, PubMed, etc.), and the literature included in well-known academic conferences (such as NeurIPS, ICML, ICLR, AAAI, ACL, and CVPR, etc.). The screening process focuses on considering factors such as the reputation of the data source, the quality standard of the included literature, the subject coverage, and the update frequency. Sources with high academic influence, a strict review mechanism, a wide subject coverage, and the ability to reflect the latest research trends in the field in a timely manner are preferentially selected to ensure the high quality and high relevance of the collected sample data, provide a reliable data basis for subsequent model training, avoid the interference of low-quality or irrelevant data, and improve the accuracy of the model in identifying cutting-edge technologies.
[0040] In step S1022, for the missing fields of the statistical meta-information in the screened sample data, mean filling or mode filling methods are used for supplementation. For numerical features (such as the number of citations, publication time, etc.), the mean value of the feature in the overall dataset is used for filling; for categorical features (such as the grouping of the number of authors, the productivity level of the first author, etc.), mode filling is used, that is, the category value with the highest frequency of occurrence of the feature is selected for filling. At the same time, a special missing marker is added to the missing fields as an auxiliary feature during model training to help the model identify and learn the data missing pattern, thereby reducing the negative impact of missing data on the model performance to a certain extent and improving the robustness of the model to incomplete data.
[0041] The statistical meta-information is standardized using the Min-Max scaling or Z-score method. Min-Max scaling linearly maps the data to the interval [0,1], making the data of different features have the same dimension, which is convenient for subsequent model processing; Z-score standardization calculates the mean and standard deviation of the data and converts the data into the form of a standard normal distribution, highlighting the distribution characteristics of the data. On this basis, the Z-score method is used to detect outliers. The values that exceed the preset threshold range (such as the absolute value of the Z-score is greater than 3) are regarded as outliers and are forced to be set to the corresponding boundary values (such as setting the values greater than the threshold to the threshold and the values less than the threshold to the negative threshold), effectively avoiding excessive interference of outliers on the model training process and ensuring the stability and accuracy of model training.
[0042] In step S1023, the publication time in the sample data is converted into a relative timestamp of the offset from the current time. The specific operation is to calculate the time difference between the publication time of each document and the current system time (or the set reference time point) and convert it into a relative timestamp in units of days, months, or years. This conversion method transforms the absolute time information into a numerical feature with a clear time span meaning, enabling the model to more intuitively capture the timeliness information of the documents, highlighting the potential advantages of recent documents compared to early documents in frontier identification. At the same time, it is also beneficial to eliminate the time bias of data in different time ranges during model training, improve the model's learning effect on time evolution features, and better grasp the dynamic trend of the development of frontier technologies over time.
[0043] An innovative frontier index is introduced to comprehensively quantify and evaluate the frontier nature of technical documents by combining the citation intensity and the contribution level of team members. The citation intensity is measured by the ratio of the number of citations to the publication time, which can effectively eliminate the influence of time factors on the number of citations and make the documents in different periods comparable; the contribution level of team members is quantified based on factors such as the number of authors, the productivity of the first author, and the citation frequency, reflecting the diversity of the research team's composition and the potential for interdisciplinary cooperation. According to the size of the frontier index, a binary classification label is added to each sample data to clearly identify whether the scientific and technological document belongs to the category of frontier technologies, thus constructing an accurate and standardized first training sample set to provide a clear supervision signal for model training.
[0044] In some embodiments, adding a label indicating whether the scientific and technological document belongs to the frontier technology in combination with the frontier index mark of each sample data includes step S1024 and step S1025: Step S1024: Calculate the frontier index of the scientific and technological document. The calculation formula is: ; where f represents the frontier index, p represents the number of citations, t represents the publication time, It represents the contribution ratio of the i-th author of a scientific and technological literature, and N represents the number of authors of the scientific and technological literature.
[0045] Step S1025: Set the frontier index thresholds for multiple technical fields respectively, and mark the scientific and technological literatures with a frontier index higher than the corresponding frontier index thresholds as frontier technologies.
[0046] In steps S1024 and S1025, the frontier index combines citation intensity and the contribution level of team members, that is, influence and innovation potential, to comprehensively measure the frontier nature of a paper. Compared with traditional citation indicators such as total citations or average annual citations, the frontier index can more accurately identify interdisciplinary innovation achievements by introducing team structure information. By setting thresholds for multiple technical fields, it is possible to distinguish whether the corresponding technology belongs to a frontier technology.
[0047] In step S103, the statistical meta-information encoder is constructed based on a multi-layer perceptron (MLP), and receives the preprocessed statistical meta-information as input. Through multiple fully connected layers and a non-linear activation function (such as Tanh), the MLP performs layer-by-layer non-linear transformation and feature extraction on the input statistical features, and finally maps the low-dimensional statistical meta-information to a high-dimensional space to generate a high-dimensional meta-information vector, fully mining the potential information and complex associations in the statistical features, so that it can be effectively fused with the text semantic information.
[0048] The title and abstract encoder relies on a preset embedded large language model, such as models that perform well in semantic tasks like NV_Embed_V2 and bge-en-icl, and inputs the title and abstract text data of the literature into it. Relying on the rich language knowledge and semantic rules learned from a large amount of text, the pre-trained large language model deeply encodes the text through a multi-layer Transformer architecture, extracts the semantic embedding vector of the text, and accurately captures the core semantic information contained in the literature title and abstract, including deep semantic features such as professional terms, research topics, and concept associations, providing a high-quality text semantic representation for subsequent fusion. After the title and abstract encoder extracts the semantic embedding vector from the vectorized text data information based on the preset embedded large language model, it also includes dimensionality reduction mapping of the semantic embedding vector.
[0049] The potential cross-attention module, as the core interaction component of the model, receives the meta-information high-dimensional vector from the statistical meta-information encoder and the semantic embedding vector from the title abstract encoder as inputs. Inside the module, the cross-attention mechanism is applied. The statistical feature vector is used as the query vector (Query), and the text semantic embedding vectors are used as the key vector (Key) and value vector (Value) respectively to calculate the correlation weights between the statistical features and the text semantic features. Through the scaled dot-product attention operation, the information of the two modalities is dynamically fused, enabling the model to automatically focus on the key features related to the emerging technology identification task in the latent space, generating a more comprehensive and in-depth literature feature vector, achieving the deep coupling and two-way enhancement of statistical features and text semantics, and overcoming the problem of insufficient information utilization caused by simple feature concatenation in traditional methods.
[0050] In some embodiments, the potential cross-attention module performs embedding processing on the semantic embedding vector to obtain the key-value vector, performs embedding processing on the meta-information high-dimensional vector to obtain the query vector. After fusing the key-value vector and the query vector based on the cross-attention mechanism, it is processed by a feed-forward neural network and a normalization layer to obtain the literature feature vector.
[0051] The fully connected layer receives the literature feature vector output by the potential cross-attention module. After further processing and transformation by the fully connected neural network, it maps it to a low-dimensional space suitable for the emerging technology classification task, and finally outputs the probability value of the prediction result indicating whether the literature belongs to the emerging technology, completing the binary classification prediction task.
[0052] During the model training and optimization process, the initial neural network model is comprehensively trained using the labeled data in the first training sample set. During the training process, based on the deviation between the prediction result output by the model and the actual label of the sample data, a loss function is constructed. Usually, the cross-entropy loss function is selected to quantify the difference between the prediction result and the true label. Through the backpropagation algorithm, the loss signal is propagated layer by layer back to each parameter of the model, and optimization algorithms such as the gradient descent method are used to iteratively update the model parameters, including the MLP weights of the statistical meta-information encoder, the fine-tuning parameters of the large language model in the title abstract encoder, the attention weight parameters of the potential cross-attention module, and the connection weights of the fully connected layer, etc. After continuous optimization through multiple training epochs, the model gradually learns how to accurately extract key features from statistical meta-information and text semantic information and accurately judge whether the literature belongs to the emerging technology. Finally, a well-trained emerging technology identification model is obtained. This model has the ability to efficiently and accurately identify emerging technology literature in practical applications, providing strong intelligence support and decision-making basis for researchers, institutions, etc., helping to grasp the forefront dynamics of scientific and technological development and optimize the layout of scientific research resources.
[0053] In some embodiments, the embedded large language model is obtained by pre-training based on the Transformer natural language processing model, and the pre-training step includes steps S201-S202: Step S201: Collect sample scientific and technological text data information for multiple technical fields, remove noise, invalid content and sensitive words, perform word segmentation, remove stop words and extract stems, and construct a second training sample set.
[0054] Sample scientific and technological text data information is collected for multiple technical fields. Text data such as scientific and technological literature, academic papers, and research reports covering many disciplines such as computer science, physics, biology, chemistry, and engineering are included. Ensure that the data sources are extensive, including academic databases (such as arXiv, PubMed), academic conference proceedings, and research institution websites. At the same time, ensure that the time span of the data is long enough to cover the technological development at different stages, so as to fully reflect the language characteristics and evolution laws of various technical fields.
[0055] Remove noise, invalid content, and sensitive words. Use natural language processing tools and regular expressions to filter out noise information such as garbled characters, special characters, and irrelevant tags in the text. Identify and remove invalid paragraphs that are not related to technical content, such as the copyright statement and acknowledgments section of the paper. Block or replace sensitive words to ensure the legitimacy and security of the data.
[0056] Perform word segmentation, stop word removal, and stemming. Use word segmentation tools that are suitable for the language characteristics of the corresponding technical field to segment the text into vocabulary units. Remove stop words such as "the" and "and" in English and "的" and "和" in Chinese. Use stemming algorithms to restore vocabulary to stem form, reduce vocabulary variant forms, and reduce feature dimensions.
[0057] The cleaned and preprocessed text data is organized into the second training sample set. Each sample is stored in the form of a word sequence after word segmentation, and the corresponding technical field and other meta-information are annotated. Ensure the diversity and representativeness of the sample set so that it can cover the language expressions and professional terms in different technical fields.
[0058] Step S202: Based on the second training sample set, the embedded large language model performs masked language modeling or next sentence prediction tasks for pre-training, and updates the parameters of the embedded large language model.
[0059] Exemplarily, based on the second training sample set, the embedded large language model performs a masked language modeling task. During training, the model randomly masks a portion of the vocabulary in the input sequence and then attempts to predict these masked words based on the context. Specifically, the encoder of the model encodes the masked input sequence to generate a context representation for each position; the decoder uses these representations to predict the words at the masked positions. In this way, the model learns the dependencies and semantic information between words.
[0060] When performing the masked language modeling task, the difference between the prediction result and the true word is calculated, and the gradient of the loss function with respect to the model parameters is calculated using the backpropagation algorithm. The model parameters are updated using an optimization algorithm (such as Adam) to gradually reduce the prediction error. After multiple rounds of iterative training, the model learns general language knowledge and semantic features from large-scale text data, forming the ability to deeply understand natural language, laying a foundation for subsequent applications in the frontier technology identification task.
[0061] In some embodiments, the embedded large language model employs models such as BERT, RoBERTa, GPT, Deepseek, NV_Embed_V2, and bge-en-icl.
[0062] On the other hand, the present invention also provides a multi-dimensional based frontier technology identification method, the method comprising steps S301 to S303: Step S301: Obtain the statistical meta-information and text data information of the scientific and technological literature to be analyzed. The statistical meta-information includes publication time, citation quantity, times cited, number of authors, productivity of the first author, and times cited. The text data information includes the title and abstract.
[0063] Step S302: Perform normalization processing on the statistical meta-information.
[0064] Step S303: Input the normalized statistical meta-information and text data information into the frontier technology identification model in the above-mentioned multi-dimensional based frontier technology identification model training method of steps S101 to S103 to output the identification result of whether the scientific and technological literature to be analyzed belongs to frontier technology.
[0065] On the other hand, the present invention also provides a multi-dimensional based frontier technology identification device, including a processor, a memory, and a computer program / instructions stored on the memory. The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the device implements the steps of the above method.
[0066] On the other hand, the present invention also provides a computer-readable storage medium, on which computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0067] On the other hand, the present invention also provides a computer program product, including a computer program / instructions, which implement the steps of the above method when executed by a processor.
[0068] The present invention will be described below in conjunction with a specific embodiment: This embodiment proposes an innovative frontier recognition framework that combines large language models (LLMs) with multi-dimensional statistical data. This integration significantly enhances the ability of large language models in frontier knowledge discovery through statistical data guidance, showing obvious advantages compared with traditional methods. Specifically, this framework utilizes the advanced natural language understanding ability of large language models, and at the same time realizes in-depth literature analysis through multi-variate statistical features, establishing a new paradigm for efficiently identifying innovative research directions in a large number of academic resources.
[0069] The structure of the present invention is shown in FIG. 2. The model mainly consists of four parts: a statistical meta-information encoder, a text encoder, a latent cross-attention module, and an output embedding for downstream tasks. The statistical meta-information encoder performs non-linear transformation and dimension alignment on the basic statistical features in the original data through a multi-layer perceptron, converting the structured statistical information into a distributed representation compatible with semantic embeddings. The text encoder is constructed based on a pre-trained language model, capturing the deep semantic information of the text through the Transformer architecture. Its output contains fine-grained representations at the token level, and a global semantic representation can be obtained through a specific pooling operation.
[0070] As the core innovation module, the latent cross-attention module uses a cross-modal attention mechanism to achieve dynamic fusion of statistical features and semantic embeddings. This module takes the text embedding as the Key-Value pair and the statistical encoding as the Query, calculating the correlation weights between features through scaled dot-product attention. This design not only retains the robustness of statistical features but also inherits the semantic understanding ability of the pre-trained model, showing significant advantages especially in dealing with domain-specific terms or data-sparse scenarios. The output embedding module is customized according to specific application scenarios. The entire model adopts an end-to-end training strategy, optimizing the statistical encoding and semantic understanding objectives jointly through a cross-entropy loss function. The model also supports balancing computational efficiency and accuracy by adjusting parameters such as the number of layers and the number of attention heads of the latent cross-attention module to meet the deployment requirements in different hardware environments.
[0071] In this module, the input meta-information comes from the scientific and technological literature data crawled from the Semantic Scholar scientific research literature database. The data of each piece of literature includes but is not limited to the following aspects: Title: The title text of the literature, used to reflect the theme and research field of the literature.
[0072] Abstract: Briefly summarizes the research content of the literature, usually used to summarize the core viewpoints of the literature.
[0073] Publication time: The publication time of the literature, which helps analyze the timeliness of the research.
[0074] Number of references: The number of other references cited in the literature, reflecting the relevance and influence of the literature.
[0075] Citation frequency: The number of times the literature is cited by other literatures, which is an important indicator to measure the influence of the literature.
[0076] Number of authors: The number of authors of the literature, reflecting the scale of the research team.
[0077] Productivity of the first author: The scientific research output of the first author, which can reflect their influence in this field.
[0078] Citation frequency of the first author: The total number of citations of the first author's literature, reflecting the influence of their academic contributions.
[0079] These meta-informations are preprocessed and standardized appropriately to ensure that each piece of data can participate in the fusion correctly in subsequent calculations. Specifically, during the preprocessing of meta-informations, first, the data with multiple features being None was cleaned, and then the data from top conferences including NeurIPS, ICML, ICLR, AAAI, ACL, and CVPR were selected. Relatively speaking, they have a certain degree of frontier nature, which is beneficial to model learning.
[0080] In this embodiment, in order to more comprehensively measure the innovation and frontier nature of the literature, the Frontier Index is introduced to label the training data. This index comprehensively considers the citation times, publication time, and the diversity of the author team of the literature. The Frontier Index can effectively identify potential frontier research fields.
[0081] The calculation formula of the Frontier Index is: ; where f represents the Frontier Index, p represents the citation times, t represents the publication time, represents the contribution ratio of the i-th author of the scientific and technological literature, and N represents the number of authors of the scientific and technological literature; is the ratio of the citation times to the publication time, which can eliminate the influence of time on the citation times. The log operation is used to compress the dynamic range of the data and avoid the excessive influence of extreme values on the result. The author heterogeneity coefficient reflects the diversity of the research team, represents the contribution ratio of the i-th author, It represents the concentration of author contributions. The author concentration is inversely related to diversity. Literature with higher author diversity usually represents the possibility of interdisciplinary cooperation and may thus be more innovative. The frontier index f provides a comprehensive measurement tool to help researchers identify the most innovative potential frontier research directions from a large number of literatures. By using the frontier index to annotate the data, the effect is better than that of the traditional average citation frequency annotation.
[0082] Frontier index thresholds are set separately for multiple technical fields, and scientific and technological literatures with a frontier index higher than the corresponding frontier index threshold are marked as frontier technologies. In this way, labels are added to the sample data to guide the training of the model.
[0083] Furthermore, the statistical meta - information encoder is a key module in the present invention for processing the statistical features of scientific and technological literatures. The main task of this module is to transform the multi - dimensional meta - information from scientific and technological literatures into high - dimensional embeddings for multi - modal fusion together with the text information. The statistical meta - information includes data such as the publication time, citation quantity, cited frequency, number of authors, productivity and cited frequency of the first author of the literature. Through appropriate transformation and encoding, these information form high - dimensional feature vectors for subsequent calculations.
[0084] To transform these meta - information into high - dimensional vectors, a multi - layer perceptron (MLP) is used for encoding. First, all input numerical data (such as publication time, citation times, number of authors, etc.) need to be normalized (BatchNorm) to have a unified scale. Specifically, each piece of data has multiple dimensions, and each dimension expresses different features. Normalization is performed according to the feature dimension. This step can be completed through the following formula: ; x is all the values of a certain feature, is the mean of this feature, is the standard deviation of this feature. After all statistical features and text features are normalized and encoded, they will be further processed by a multi - layer perceptron (MLP) to be transformed into high - dimensional embeddings. The basic structure of the MLP model consists of multiple fully - connected layers, and the transformation of features is achieved through a non - linear activation function (Tanh). Specifically, the processing process of the multi - layer perceptron is represented by the following steps: Input layer: Concatenate all normalized numerical features (publication time, citation times, number of authors, etc.) and text features (title, abstract) into a vector.
[0085] Hidden layer: This vector undergoes non - linear transformation through the hidden layer in the MLP to form an intermediate feature representation.
[0086] Output layer: Finally, map the output of the hidden layer to the target high-dimensional embedding space through the output layer of the MLP.
[0087] The output of the MLP is expressed as: ; is the i-th data The embedding after dimensionality-increasing feature mapping. The statistical meta-information encoder can effectively represent the influence, timeliness, and innovation of the author team of the literature by transforming multi-dimensional meta-information into high-dimensional embeddings.
[0088] Furthermore, the text encoder is the core module in this embodiment for processing scientific literature text data. The main task of this module is to transform text information such as the title and abstract of the literature into high-dimensional embeddings for fusion processing with statistical meta-information, so as to achieve more accurate identification of cutting-edge research directions. The text encoder uses natural language processing (NLP) techniques, especially pre-trained large language models, to capture semantic information in the literature.
[0089] To better capture the semantic information of the text, referring to the MTEB list, the text encoder selects large language models (NV_Embed_V2 and bge-en-icl) that perform well in various semantic tasks. These models are trained on large-scale text datasets, can learn rich language rules and knowledge, and are at the leading level in processing various NLP tasks. Therefore, when processing scientific research literature, they can effectively capture domain-related terms and deep semantics.
[0090] The title and abstract of the literature are used as inputs. After being processed by the pre-trained language model, a high-dimensional semantic embedding vector is output, and the specific processing process is as follows: Input layer: The title of the literature and the abstract are fed into the pre-trained language model for encoding.
[0091] Encoding process: The pre-trained language model processes the input text through multiple Transformer layers to extract the deep semantic information in the text.
[0092] Output layer: Finally, the model outputs a fixed-length high-dimensional vector representing the semantic features of the literature.
[0093] This process can be represented by the following formula: ; The text encoder can effectively capture the semantic information in the text by converting the title and abstract of the literature into high-dimensional embeddings. These semantic features play a crucial role in the identification process of cutting-edge research. After being combined with the statistical meta-information encoder, the text encoder provides rich semantic features, providing strong support for the identification of the cutting-edge research direction of the literature.
[0094] Furthermore, the potential cross-attention module aims to fuse text information and statistical meta-information through deep learning methods to improve the accuracy of cutting-edge research direction identification. This module adopts the cross-attention mechanism to capture the complex associations between multi-modal data, thereby generating more representative literature feature representations.
[0095] Before entering the potential cross-attention module, the high-dimensional feature representations from the text encoder and the statistical meta-information encoder have been completed. Specifically: Text features Generated by the text encoder, representing the semantic information of the literature.
[0096] Statistical features Generated by the statistical meta-information encoder, representing the metadata of the literature (such as the number of citations, publication time, etc.) These two sets of features will be passed as inputs into the potential cross-attention module, and a comprehensive literature feature representation will be generated through deep fusion .
[0097] The core innovation of the potential cross-attention module lies in the introduction of the cross-attention mechanism, enabling the model to perform two-way interaction between text information and statistical information in the latent space. Specifically, the text information and statistical information will be aligned in a specific "latent space" to effectively capture the relationship between the two.
[0098] The basic process of cross-attention is as follows: Suppose there are text vectors and statistical vectors , first map them to transform them into the same latent space. This process is represented by the following formula: , ; , ; Among them, and are weight matrices learned during the training process, used to generate queries and keys respectively. Next, based on the mappings of these queries and keys, calculate the cross-attention weights: ; ; is the calculated attention weight, is the obtained attention fusion result, represents the value vector, which is obtained by multiplying the text vector with the trainable parameter matrix The cross-attention calculation adjusts the value vector according to the obtained attention weight to obtain the fused representation, so as to better fuse the text and statistical information.
[0099] After the cross-attention calculation, the obtained feature is used as the intermediate feature. Next, it is further fused and feature-learned through a feed-forward neural network (FFN). The specific fusion process is as follows: ; Through the feed-forward neural network, the model further combines the features from the text and statistical information to generate the final comprehensive representation of the literature , and normalizes the result (LayerNorm).
[0100] Finally, what the potential cross-attention module outputs is a new high-dimensional feature vector , which combines the deep semantics and statistical associations of the text and statistical information. This feature vector will be used as input in subsequent tasks to help identify potential emerging research directions.
[0101] Through this cross-attention mechanism, the potential cross-attention module realizes the deep fusion of multi-modal information and can effectively improve the ability to identify the emerging nature of the literature. Compared with the traditional simple concatenation method, the potential cross-attention module significantly improves the representation ability and accuracy of the model.
[0102] By introducing the cross-attention mechanism, the potential cross-attention module realizes the two-way interaction and fusion of text information and statistical meta-information in the latent space. This method not only enhances the model's understanding of multi-modal information but also improves the accuracy of identifying emerging research directions. By deeply fusing text and statistical information, the potential cross-attention module provides an efficient way of feature learning and strongly supports the identification of emerging research directions in the literature.
[0103] Correspondingly to the above method, the present invention further provides a device / system, which includes a computer device. The computer device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device / system implements the steps of the method as described above.
[0104] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the foregoing edge computing server deployment method. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the art.
[0105] This embodiment proposes an innovative frontier index calculation method for quantifying the innovation and frontier nature of scientific research literature. Compared with the single indicators of traditional citation times and publication time, this embodiment provides a more comprehensive and accurate frontier evaluation criterion by comprehensively considering multi-dimensional features such as the citation times, publication time, and diversity of the author team of the literature. This method significantly improves the ability to identify potential frontier research fields. Especially for frontier fields that have not been widely cited, it can provide more accurate predictions. The potential cross-attention module is used to process and fuse multi-modal data (such as text information and statistical meta-information). The module adopts an advanced cross-attention mechanism, which can achieve deep interaction between text information and statistical information in the latent space and generate a high-quality literature feature representation. Through the combination of the self-attention mechanism and the cross-attention mechanism, the potential cross-attention module can efficiently capture the association between text and statistical information and improve the accuracy of identifying the frontier research direction of the literature. Compared with the traditional simple splicing method, the potential cross-attention module significantly enhances the performance of the frontier field identification model through deep feature fusion and optimization.
[0106] In summary, the multi-dimensional based frontier technology identification model training method, identification method and device of the present invention introduce multi-dimensional statistical meta-information and text data information to construct a training sample set, and fuse the citation intensity and the contribution level of team members to quantify the frontier nature of technical literature for accurate tagging. The model introduces a multi-layer perceptron to extract the high-dimensional features of statistical meta-information, introduces a pre-trained embedded large language model to extract the semantic features of the titles and abstracts of scientific and technological documents, and performs frontier technology identification after feature fusion through a potential cross self-attention module, realizing the deep integration of multi-modal information, enabling the model to comprehensively and accurately grasp the features of scientific and technological literature, and significantly improving the accuracy and timeliness of frontier technology identification during the application process.
[0107] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or a communication link.
[0108] It should be clear that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0109] In the present invention, features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0110] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and variations can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A training method for a multi-dimensional cutting-edge technology identification model, characterized in that, The method includes the following steps: Obtain multiple pieces of sample data, each piece of sample data including statistical meta-information and text data information of scientific and technological literature. The statistical meta-information includes publication time, citation quantity, cited frequency, number of authors, productivity of the first author, and cited frequency. The text data information includes title and abstract; Perform preprocessing on each piece of the sample data. The preprocessing includes data screening, data cleaning and standardization of the statistical meta-information, and adding a label indicating whether the scientific and technological literature belongs to a frontier technology by combining the frontier index mark of each piece of sample data to form a first training sample set; the frontier index quantifies the frontier nature of technical literature by combining citation intensity and the contribution level of team members; Obtain an initial neural network model, including a statistical meta-information encoder, a title and abstract encoder, a potential cross-attention module, and a fully connected layer. The statistical meta-information encoder converts the preprocessed statistical meta-information into a high-dimensional meta-information vector based on a multi-layer perceptron. The title and abstract encoder extracts a semantic embedding vector from the vectorized text data information based on a preset embedded large language model. The potential cross-attention module takes the high-dimensional meta-information vector and the semantic embedding vector as inputs and outputs a literature feature vector. The literature feature vector is input into the fully connected layer to output a prediction result indicating whether it belongs to a frontier technology; use the first training sample set to train the initial neural network model, and construct a loss to update the model parameters based on the deviation between the prediction result and the label to obtain a frontier technology recognition model.
2. The method for training a cutting-edge technology identification model based on multiple dimensions according to claim 1, wherein Perform preprocessing on each piece of the sample data, including: Screen the sample data according to a preset number of scientific and technological literature sources; Fill the missing fields of the statistical meta-information in the screened sample data with the mean or the mode and add a missing mark for auxiliary identification; standardize the statistical meta-information using Min-Max scaling or Z-score, detect outliers through Z-score, mark the values exceeding the threshold range as outliers and replace them with boundary values; Convert the publication time to a relative timestamp of the offset from the current time.
3. The method for training a cutting-edge technology identification model based on multiple dimensions according to claim 1, wherein Adding a label indicating whether the scientific and technological literature belongs to a frontier technology by combining the frontier index mark of each piece of sample data includes: Calculate the frontier index of the scientific and technological literature, and the calculation formula is: ; Among them, f represents the frontier index, p represents the citation times, t represents the publication time, represents the contribution ratio of the i-th author of the scientific and technological literature, and N represents the number of authors of the scientific and technological literature; Set a frontier index threshold for each of multiple technical fields respectively, and mark the scientific and technological literature with a frontier index higher than the corresponding frontier index threshold as a frontier technology.
4. The training method of the cutting-edge technology identification model based on multiple dimensions according to claim 1, wherein The embedded large language model is pre-trained using a Transformer-based natural language processing model. The pre-training steps include: Collect sample scientific and technological text data information for multiple technical fields, remove noise, invalid content, and sensitive words, perform word segmentation, stop word removal, and stemming extraction, and construct a second training sample set; Based on the second training sample set, the embedded large language model performs masked language modeling or next sentence prediction tasks for pre-training to update the parameters of the embedded large language model.
5. The method for training a cutting-edge technology identification model based on multiple dimensions according to claim 4, wherein The embedded large language model adopts BERT, RoBERTa, GPT, Deepseek, NV_Embed_V2 and bge-en-icl models.
6. The method for training a cutting-edge technology identification model based on multiple dimensions according to claim 4, wherein The potential cross-attention module performs embedding processing on the semantic embedding vector to obtain a key-value vector, performs embedding processing on the meta-information high-dimensional vector to obtain a query vector, and after fusing the key-value vector and the query vector based on the cross-attention mechanism, obtains the literature feature vector after processing by a feed-forward neural network and a normalization layer; After the title abstract encoder extracts the semantic embedding vector from the vectorized text data information based on a preset embedded large language model, it further includes performing dimensionality reduction mapping on the semantic embedding vector.
7. A multi-dimensional based cutting-edge technology identification method, characterized in that, The method includes: Obtaining the statistical meta-information and text data information of the scientific and technological literature to be analyzed, where the statistical meta-information includes publication time, citation quantity, times cited, number of authors, productivity of the first author, and times cited, and the text data information includes title and abstract; Performing standardization processing on the statistical meta-information; Inputting the standardized statistical meta-information and the text data information into the frontier technology recognition model in the multi-dimensional frontier technology recognition model training method according to any one of claims 1 to 6 to output the recognition result of whether the scientific and technological literature to be analyzed belongs to the frontier technology.
8. A multi-dimensional-based cutting-edge technology identification device, comprising a processor, a memory, and computer programs / instructions stored on the memory, characterized in that, The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the device implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Technical research frontier index calculation method
CN108595559A
Front-edge-technology-oriented calculation method for the front-edge index of a scientific research institution
CN108629489A
Model training method applied to text recognition and text recognition method and device
CN114625874A
Academic literature recommendation method fusing title and abstract semantic relation
CN114626369A
Method for analyzing influence index of literature
CN115357731A