Paper examination method and equipment based on multi-dimensional data fusion
Through the multi-dimensional data fusion method, combined with research records and paper text data, an interactive attention mechanism model is constructed, which solves the problems of high misjudgment rate of paper review and high resource occupancy rate in the existing technology, and achieves more accurate and efficient paper review.
Patent Information
- Application Number
- CN202510747255.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing automated review methods for papers have high misjudgment rates when facing emerging research directions, and the storage cost and resource utilization rate of blockchain technology are high, making it difficult to effectively identify whether the paper is temporarily pieced together, resulting in unreasonable review scores.
Through the multi-dimensional data fusion method, natural language processing technology and interactive attention mechanism are used, combined with research record data and paper text data, time dimension scores, spatial variables, text correlation coefficients and content quality feature vectors are obtained, and multi-dimensional data fusion paper review model is constructed, and review scores are output.
It improves the rationality and effectiveness of paper review, reduces the equipment resource occupancy rate and storage cost, reduces the misjudgment rate, and ensures the accuracy of the review score.
Smart Images

Figure CN120257978A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of electronic digital data processing, and in particular, to a paper review method and device based on multi-dimensional data fusion. Background Art
[0002] Paper review refers to the process of examining and evaluating aspects such as the content and structure of academic papers. Its purpose is to ensure the academic nature, standardization, and integrity of papers, and improve the quality and level of papers. The traditional review and assessment method is for experts to conduct reviews and give paper review scores. However, this way of manually reviewing papers not only takes a long time, but also, due to over-reliance on templated requirements, may cause paper authors to sacrifice the coherence of the content to meet format specifications, and the subjectivity of manual review makes it difficult to effectively identify the authenticity of the research process, difficult to trace the real behaviors during the research cycle, and also difficult to determine whether the paper content is derived from literature piecing or real research accumulation. Therefore, the current automated implementation of paper review has become a research focus.
[0003] Currently, in the existing automated paper review methods, there is the method of AI paper quality profiling, that is, the AI collaborative review system predicts the risk of paper withdrawal by training a "paper quality profiling" model and combining historical review data. There is also the use of blockchain technology to deposit and trace the experimental data related to papers, deposit the experimental data, modification records, etc. on the chain to ensure the data cannot be tampered with. Or automatically identify the association between the author and the funder through smart contracts to prevent situations such as conflicts of interest.
[0004] However, the method of AI paper quality profiling relies on historical data to train the "paper quality profiling" model. If there is too little or no historical data available for reference in emerging research directions, there is a high misjudgment rate. Although blockchain technology can ensure that the data cannot be tampered with and can deposit the experimental data and modification records on the chain, it faces high occupancy of device resources, high storage costs, and technical thresholds.
[0005] Therefore, there is an urgent need to design a method that can improve the rationality and effectiveness of automated paper review, and can reduce the occupancy rate of device resources on the basis of realizing the judgment of whether a paper is pieced together temporarily. Summary of the Invention
[0006] In view of this, the embodiments of this application provide a paper review method and device based on multi-dimensional data fusion to eliminate or improve one or more defects existing in the prior art.
[0007] One aspect of this application provides a paper review method based on multi-dimensional data fusion, including: From a software system for storing research record data and check-in information periodically submitted by users, extract research record data that matches the current target paper as target record data, and based on the target record data and the paper text data corresponding to the target paper, respectively obtain the time dimension score for representing the time distribution of the research of the target paper, the spatial variable for representing the spatial distribution of the research of the paper, and the text correlation coefficient for representing the similarity between the paper text data and the target record data; And, based on natural language processing technology, determine the content quality feature vector for representing the content quality of the target paper according to the paper text data, and determine the text quantity index for representing the proportion of the quantity dimension information in the target paper according to the paper text data; Input the time dimension score, the text correlation coefficient, the text quantity index, the spatial variable, and the content quality feature vector corresponding to the target paper into a multi-dimensional data fusion paper review model based on an interactive attention mechanism, so that the paper review model outputs the paper review score corresponding to the target paper.
[0008] In some embodiments of the present application, in the software system for storing research record data and check-in information periodically submitted by users, extracting research record data that matches the current target paper as target record data includes: In the software system for storing research record data and check-in information periodically submitted by users, extract each research record data periodically submitted by the user who is the author of the target paper; Respectively extract the paper text data corresponding to the target paper and the keywords corresponding to each of the research record data, and calculate the frequency and inverse document frequency of each keyword appearing in the paper text data or the research record data where it is located; According to the frequency and inverse document frequency corresponding to each keyword, use the term frequency-inverse document frequency algorithm to find the research record data with keyword matching between the target paper in each research record data as the target record data.
[0009] In some embodiments of the present application, the obtaining, according to the target record data and the paper text data corresponding to the target paper, the time dimension score for representing the time distribution of the research of the target paper, the spatial variable for representing the spatial distribution of the research of the paper, and the text correlation coefficient for representing the similarity between the paper text data and the target record data respectively includes: Extract the submission time corresponding to the target record data from the software system to obtain the submission time period corresponding to all the target record data; determine the time dimension score corresponding to the target paper for representing the time distribution of the paper research according to the submission time period and a preset empirical threshold; Obtain the spatial variable corresponding to the target paper for representing the spatial distribution of the paper research according to the target record data and the paper text data corresponding to the target paper; And obtain the text correlation coefficient corresponding to the target paper for representing the similarity between the paper text data and the target record data according to the target record data and the paper text data corresponding to the target paper.
[0010] In some embodiments of the present application, the obtaining the spatial variable corresponding to the target paper for representing the spatial distribution of the paper research according to the target record data and the paper text data corresponding to the target paper includes: Obtain the check-in information corresponding to each of the target record data submitted by the user who is the author of the target paper in the software system, and respectively determine the geographical location information corresponding to each of the check-in information according to the geographic information system; Perform periodic encoding on each of the geographical location information to obtain the spatial variable corresponding to the target paper for representing the spatial distribution of the paper research.
[0011] In some embodiments of the present application, the obtaining the text correlation coefficient corresponding to the target paper for representing the similarity between the paper text data and the target record data according to the target record data and the paper text data corresponding to the target paper includes: Generate the abstract text data corresponding to the target record data based on the large language model; Input the abstract text data and the paper text data into the BERT model respectively, so that the BERT model outputs the abstract feature vector corresponding to the abstract text data and the paper feature vector corresponding to the paper text data respectively; Calculate the similarity between the abstract feature vector and the paper feature vector by using cosine similarity as the text correlation coefficient corresponding to the target paper for representing the similarity between the paper text data and the target record data.
[0012] In some embodiments of the present application, the determining the content quality feature vector corresponding to the target paper for representing the content quality of the paper according to the paper text data based on the natural language processing technology includes: Preprocess the paper text data; wherein, the preprocessing includes: removing noise, clause segmentation, and word segmentation; Based on natural language processing technology, corresponding to a tokenizer, obtain a word embedding vector sequence corresponding to the preprocessed paper text data; Perform a linear mapping on the hidden layer feature vectors corresponding to the word embedding vector sequence to obtain a content quality feature vector corresponding to the target paper for representing the quality of the paper content.
[0013] In some embodiments of the present application, the determining the text quantity index corresponding to the target paper for representing the proportion of quantity dimension information in the paper according to the paper text data includes: Obtain the data volume, number of experiments, and number of cited references corresponding to the paper text data; Perform standardization processing on the data volume, number of experiments, and number of cited references respectively to obtain a first standard value corresponding to the data volume, a second standard value corresponding to the number of experiments, and a third standard value corresponding to the number of cited references; According to the first standard value, the second standard value, and the second standard value and their respective corresponding weights, calculate to obtain an original quantity index corresponding to the paper text data; Perform a linear transformation on the original quantity index to convert it to the interval [0, 10] to obtain a text quantity index corresponding to the target paper for representing the proportion of quantity dimension information in the paper.
[0014] In some embodiments of the present application, the multi-dimensional data fusion paper review model based on an interactive attention mechanism includes: A scalar embedding layer for performing scalar embedding on the time dimension score, the text correlation coefficient, and the text quantity index corresponding to the target paper to obtain a time feature vector corresponding to the time dimension score, a text correlation feature vector corresponding to the text correlation coefficient, and a text quantity feature vector corresponding to the text quantity index; A vector projection layer for performing vector projection on the spatial variable and the content quality feature vector corresponding to the target paper to obtain a projected spatial variable corresponding to the spatial variable and a projected content quality feature vector corresponding to the content quality feature vector; An input sequence construction layer for constructing a corresponding input sequence according to the time feature vector, the text correlation feature vector, the text quantity feature vector, the projected spatial variable, and the projected content quality feature vector; A spatial and content cross-attention layer, which is used to take the projected content quality feature vector in the input sequence as a query vector, and take the projected spatial variable as a key-value pair vector to perform cross-attention calculation to obtain a corresponding spatial and content fusion feature vector; A quantity and quality cross-attention layer, which is used to splice the projected content quality feature vector in the input sequence and the spatial and content fusion feature vector to obtain a corresponding first spliced vector; take the first spliced vector as a query vector, and take the text quantity feature vector in the input sequence as a key-value pair vector to perform cross-attention calculation to obtain a corresponding quantity and quality fusion feature vector; A vector splicing layer, which is used to splice the projected content quality feature vector, the projected spatial variable, and the quantity and quality fusion feature vector in the input sequence to obtain a corresponding second spliced vector; A dynamic fusion layer, which is used to dynamically fuse the second spliced vector with the text correlation feature vector in the input sequence as a modulation factor; A three-dimensional attention routing layer, which is used to obtain the target output representation corresponding to the target paper according to the self-interaction result data between the time feature vector, the text correlation feature vector, and the text quantity feature vector, the self-interaction result data between the projected content quality feature vector and the projected spatial variable, and the second spliced vector after dynamic fusion; A linear layer, which is used to map the target output representation to the paper review score corresponding to the target paper.
[0015] Another aspect of the present application provides a paper review device based on multi-dimensional data fusion, including: A first multi-dimensional data acquisition module, which is used to extract research record data matching the current target paper from a software system for storing research record data and check-in information periodically submitted by users as target record data, and respectively obtain a time dimension score representing the time distribution of the paper research, a spatial variable representing the spatial distribution of the paper research, and a text correlation coefficient representing the similarity between the paper text data and the target record data corresponding to the target paper according to the target record data and the paper text data corresponding to the target paper; And a second multi-dimensional data acquisition module, which is used to determine a content quality feature vector representing the content quality of the target paper according to the paper text data based on natural language processing technology, and determine a text quantity index representing the proportion of quantity dimension information in the target paper according to the paper text data; The paper review scoring module is used to input the time dimension score, the text correlation coefficient, the text quantity index, the spatial variable, and the content quality feature vector corresponding to the target paper into the multi-dimensional data fusion paper review model based on the interactive attention mechanism, so that the paper review model outputs the paper review score corresponding to the target paper.
[0016] The third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the paper review method based on multi-dimensional data fusion is implemented.
[0017] The fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the paper review method based on multi-dimensional data fusion is implemented.
[0018] The fifth aspect of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the paper review method based on multi-dimensional data fusion is implemented.
[0019] The paper review method based on multi-dimensional data fusion provided by this application extracts research record data matching the current target paper from a software system used to store research record data and check-in information periodically submitted by users as target record data, and respectively obtains a time dimension score representing the time distribution of the paper research corresponding to the target paper, a spatial variable representing the spatial distribution of the paper research, and a text correlation coefficient representing the similarity between the paper text data and the target record data according to the target record data and the paper text data corresponding to the target paper; and, based on natural language processing technology, determines a content quality feature vector representing the content quality of the target paper corresponding to the target paper according to the paper text data, and determines a text quantity index representing the proportion of quantity dimension information in the target paper according to the paper text data; inputs the time dimension score, the text correlation coefficient, the text quantity index, the spatial variable, and the content quality feature vector corresponding to the target paper into a multi-dimensional data fusion paper review model based on an interactive attention mechanism, so that the paper review model outputs a paper review score corresponding to the target paper; by extracting research record data periodically submitted by users matching the current target paper from a software system used to store research record data and check-in information periodically submitted by users as target record data, and respectively obtaining the time dimension score, the spatial variable, and the text correlation coefficient corresponding to the target paper according to the target record data and the paper text data corresponding to the target paper, it can effectively trace the user's daily paper research time, research space distribution, and similarity, and can effectively improve the recognition accuracy of whether the paper content is pieced together temporarily; and, by using the time dimension score, the text correlation coefficient, the text quantity index, the spatial variable, and the content quality feature vector corresponding to the target paper as multi-dimensional data for automatic scoring of the paper review score, it can improve the rationality and effectiveness of the paper review score; that is to say, this application can improve the accuracy of judging whether a paper is pieced together temporarily without using blockchain technology, and can effectively reduce the resource occupancy rate and storage cost of the device for performing paper review and lower the technical threshold; and compared with the method of using an AI paper quality portrait, it can effectively improve the rationality and effectiveness of the paper review score and reduce the misjudgment rate.
[0020] Additional advantages, objects, and features of this application will be partly described below, and will become partly apparent to those of ordinary skill in the art after studying the following parts, or can be learned from the practice of this application. The objects and other advantages of this application can be achieved and obtained by the structures specifically pointed out in the specification and the drawings.
[0021] Those skilled in the art will understand that the objectives and advantages achievable with the present application are not limited to those specifically described above, and the above and other objectives achievable with the present application will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings described herein are used to provide a further understanding of the present application, form a part of the present application, and do not limit the present application. The components in the drawings are not drawn to scale, but are only for showing the principles of the present application. For the convenience of showing and describing some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings: Figure 1 FIG. 8 is a first schematic flowchart of a paper review method based on multi-dimensional data fusion in an embodiment of the present application.
[0023] Figure 2 FIG. 12 is a second schematic flowchart of a paper review method based on multi-dimensional data fusion in an embodiment of the present application.
[0024] Figure 3 FIG. 16 is a third schematic flowchart of a paper review method based on multi-dimensional data fusion in an embodiment of the present application.
[0025] Figure 4 FIG. 20 is a schematic architecture diagram of a multi-dimensional data fusion paper review model based on an interactive attention mechanism in an embodiment of the present application.
[0026] Figure 5 FIG. 24 is a schematic structural diagram of a paper review device based on multi-dimensional data fusion in an embodiment of the present application.
[0027] Figure 6 FIG. 28 is an application schematic diagram of a multi-dimensional data fusion paper review model based on an interactive attention mechanism in an application example of the present application.
[0028] Figure 7 FIG. 32 is a schematic usage flowchart of a postgraduate monthly report system in an application example of the present application.
[0029] Figure 8 FIG. 36 is a schematic flowchart of a student submitting monthly or annual report information in an application example of the present application.
[0030] Figure 9 FIG. 40 is a schematic flowchart of the management terminal reviewing the monthly or annual report information submitted by students in an application example of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below in conjunction with the embodiments and the drawings. Herein, the illustrative embodiments of the present application and their descriptions are used to explain the present application, but not to limit the present application.
[0032] Herein, it should also be noted that in order to avoid obscuring the present application with unnecessary details, only the structures and / or processing steps closely related to the solution according to the present application are shown in the drawings, while other details less relevant to the present application are omitted.
[0033] It should be emphasized that the term "comprising / including" as used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0034] Herein, it should also be noted that if not otherwise specified, the term "connection" as used herein can refer not only to direct connection, but also to indirect connection with an intermediate.
[0035] In the following, embodiments of the present application will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0036] It should be noted that the general processing flow of traditional paper review and assessment methods is as follows: First, the editor conducts a preliminary screening and formal review of the paper, and then the reviewers conduct multiple rounds of review, including plagiarism detection and data authenticity review. Finally, an examination of innovation and other aspects is carried out before the paper can enter the publication process. However, when faced with the serious disconnection between the content of postgraduate papers and their daily research, the following drawbacks may be exposed: (1) At the level of formal review, over-reliance on templated requirements may lead students to sacrifice the coherence of the content in order to meet the format specifications. For example, some students forcefully insert randomly pieced-together chapters into the framework to meet the journal format requirements, while traditional formal review only focuses on explicit elements such as figure numbers and reference formats, and it is difficult to identify the disconnection between data authenticity and research logic. (2) In terms of content review, the subjectivity of peer review makes it difficult to effectively identify the authenticity of the research process. When the paper is disconnected from daily research, although the single-blind or double-blind review mechanism can reduce author identity bias, it cannot trace the original scenario of data generation. For example, some experimental data may be tampered with or fabricated to fit the conclusion, and traditional review mainly relies on expert experience judgment, lacking dynamic verification means for the research process. (3) If postgraduate students fabricate the research process in their papers, the current review mechanism often only focuses on the formal declarations at the time of submission and is difficult to trace their real behaviors during the research cycle. For example, some students may overly polish data interpretations through proofreading services or even purchase ghostwriting services, but such behaviors are difficult to detect under the traditional review framework, resulting in the long-term concealment of academic misconduct risks. (4) In terms of innovation assessment. When the content of the paper is disconnected from daily research, its innovation may stem from literature piecing rather than real research accumulation. However, the current system relies on quantitative indicators such as citation counts and Altmetrics, and it is difficult to distinguish between superficial innovation and substantial breakthroughs, resulting in low-quality papers passing the formal review from time to time. Altmetrics is a method for evaluating research impact, which aims to evaluate the impact of academic achievements by quantitatively analyzing the dissemination and citation of academic achievements on platforms such as social media, news media, and policy documents. Its emergence is to address the limitations of traditional academic evaluation indicators (such as impact factors) and provide a more comprehensive, timely, and diverse academic evaluation method.
[0037] Regarding the methods of paper quality control and management, there are many mature and applied methods, and there are also some emerging methods. For example, the AI collaborative review system predicts the risk of paper withdrawal by training the "paper quality portrait" model and combining historical review data, and has achieved a high prediction accuracy in actual application. With the continuous development and improvement of the Internet system, blockchain technology is also used to store and trace the experimental data related to the paper, and the experimental data, modification records, etc. are stored on the chain to ensure that the data cannot be tampered with. Or the relationship between the author and the sponsor is automatically identified through smart contracts to prevent conflicts of interest. Therefore, in order to accurately and comprehensively control the quality of submitted papers and meet the requirements of high-level and high-quality papers, it is necessary to establish a paper submission review method that combines process management with goal management. This combination of process and goal not only ensures the coherence and scientificity of the paper from research to writing, but also improves the level of the paper from the source, reduces the output of low-quality papers due to confusion in the research process or unclear goals, and effectively promotes academic research in the direction of high quality.
[0038] In the field of paper quality control, the emerging AI collaborative review system and blockchain technology have brought many changes, but they also have certain defects. The AI collaborative review system relies on historical data to train the "paper quality portrait" model, but if there is too little or no historical data to refer to for a new research direction, there will be a high rate of misjudgment. Although blockchain technology can ensure that data cannot be tampered with and experimental data and modification records can be stored on the chain, it faces high storage costs and technical barriers. Not all research teams and editorial departments are able to bear and use it.
[0039] Compared with these, the method of combining process management and target management proposed in this application has obvious advantages. From the perspective of process management, the duration of research is traced through the time data of the journal system, and the research location is verified with the help of spatial information to supervise the research progress in all aspects to avoid temporary patchwork of research results. This is a deep control of the research process that AI and blockchain technology cannot achieve. In terms of target management, a comprehensive evaluation model covering quality and quantity dimensions is constructed, and papers are evaluated from multiple perspectives in combination with expert reviews. It can not only judge the current quality of the paper, but also pay attention to its consistency with daily research accumulation, and tap the value of long-term research precipitation. This comprehensive, in-depth and flexible method that can be adapted to various types of papers will not be limited by missing data or technical difficulties. It effectively guarantees the quality of papers from the source and provides solid support for the healthy development of academic research.
[0040] Based on this, in order to solve the problem that the existing method for the quality portrait of AI papers relies on historical data to train the "paper quality portrait" model. If there is too little or no historical data available for emerging research directions, the rationality and effectiveness of the paper review scores cannot be guaranteed. Also, in order to solve the problems faced by the blockchain method, such as high occupancy rate of device resources, high storage costs, and technical thresholds, embodiments of the present application respectively provide a paper review method based on multi-dimensional data fusion, a paper review device based on multi-dimensional data fusion for executing the paper review method based on multi-dimensional data fusion, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the accuracy of ad-hoc judgment in paper review, thereby effectively reducing the resource occupancy rate and storage costs of the device for executing paper review, and improving the rationality and effectiveness of paper review scores.
[0041] Specific details are described in detail through the following embodiments.
[0042] Based on this, embodiments of the present application provide a paper review method based on multi-dimensional data fusion that can be implemented by a paper review device based on multi-dimensional data fusion. Refer to Figure 1 , the paper review method based on multi-dimensional data fusion specifically includes the following content: Step 100: Extract research record data matching the current target paper from a software system for storing research record data and check-in information periodically submitted by users as target record data, and respectively obtain a time dimension score representing the time distribution of the paper research, a spatial variable representing the spatial distribution of the paper research, and a text correlation coefficient representing the similarity between the paper text data and the target record data corresponding to the target paper according to the target record data and the paper text data corresponding to the target paper.
[0043] The software system is used to store research record data and check-in information periodically submitted by users. The period for submitting research record data can be set according to actual application requirements. For example, it can be set to weekly, monthly, or quarterly, etc. If the period is monthly, the research record data is the monthly report. In an example of the present application, the software system can adopt a graduate monthly report system, and the specific functions will be described in detail in subsequent application examples and will not be elaborated here.
[0044] The target paper refers to the academic paper to be currently reviewed.
[0045] The time dimension score is used to represent the time distribution of the paper research, which can be represented by S; its score can be from 0 to 10. The higher the score, the more the target paper tends to be a real research result; the lower the score, the more the target paper tends to be ad-hoc.
[0046] The spatial variable is used to represent the spatial distribution of the research space of the paper and can be denoted by D; the spatial variable can be represented in a periodic coding manner.
[0047] The text correlation coefficient is used to represent the similarity between the paper text data and the target record data, can be denoted by k, and can be specifically obtained by calculating the cosine similarity.
[0048] And, step 200: Based on natural language processing technology, determine the content quality feature vector corresponding to the target paper for representing the content quality of the paper according to the paper text data, and determine the text quantity index corresponding to the target paper for representing the proportion of the quantity dimension information in the paper according to the paper text data.
[0049] In step 200, the natural language processing (NLP) technology refers to a deep learning model constructed through large-scale pre-training and self-supervised learning technologies, aiming to improve the computer's ability to understand and generate natural language.
[0050] It can be understood that the content quality feature vector can be denoted by Q; the text quantity index can be denoted by N, and the quantity dimension information can include information such as the data volume, the number of experiments, and the number of cited documents corresponding to the target paper.
[0051] Step 300: Input the time dimension score, the text correlation coefficient, the text quantity index, the spatial variable, and the content quality feature vector corresponding to the target paper into a multi-dimensional data fusion paper review model based on an interactive attention mechanism, so that the paper review model outputs the paper review score corresponding to the target paper.
[0052] Among them, the interactive attention is an advanced attention mechanism used for information exchange and enhanced feature expression between different modalities. This mechanism enables the model to exchange information and strengthen the learning of key features when processing multi-modal data, such as different types of data streams like text and image, audio and video, etc. The interactive attention is designed to address the limitation that the traditional attention mechanism may ignore cross-modal interactions when processing complex multi-modal data.
[0053] The core advantage of interactive attention is that it can establish deep connections and collaborations among different information sources (i.e., the time - dimension score S corresponding to the target paper, the text correlation coefficient k, the text quantity index N, the spatial variable D, and the content quality feature vector Q). Through this connection, the model can capture richer and more useful information among different modalities, thereby improving the performance of the overall task.
[0054] As can be seen from the above description, the paper review method based on multi - dimensional data fusion provided by the embodiments of the present application extracts the research record data submitted periodically by the user that matches the current target paper as the target record data from the software system for storing the research record data and check - in information submitted periodically by the user, and respectively obtains the time - dimension score, spatial variable, and text correlation coefficient corresponding to the target paper according to the target record data and the paper text data corresponding to the target paper, which can effectively trace the user's daily paper research time, research space distribution, and similarity, and can effectively improve the recognition accuracy of whether the paper content is pieced together temporarily; and, by using the time - dimension score S, the text correlation coefficient k, the text quantity index N, the spatial variable D, and the content quality feature vector Q corresponding to the target paper as multi - dimensional data for automatic scoring of the paper review score, it can improve the rationality and effectiveness of the paper review score; that is to say, the present application can improve the accuracy of the temporary piecing - together judgment of paper review without using blockchain technology, and thus can effectively reduce the resource occupancy rate and storage cost of the device for performing paper review and lower the technical threshold; and compared with the method of using an AI paper quality portrait, it can effectively improve the rationality and effectiveness of the paper review score and reduce the misjudgment rate.
[0055] In order to further improve the effectiveness and reliability of extracting the research record data that matches the current target paper as the target record data, in a paper review method based on multi - dimensional data fusion provided by the embodiments of the present application, refer to Figure 2 , step 100 in the paper review method based on multi - dimensional data fusion specifically includes the following content: Step 110: Extract each research record data submitted periodically by the user who is the author of the target paper from the software system for storing the research record data and check - in information submitted periodically by the user.
[0056] Step 120: Respectively extract the paper text data corresponding to the target paper and the keywords corresponding to each of the research record data, and calculate the frequency and inverse document frequency of each keyword in the paper text data or the research record data where it is located.
[0057] Step 130: Based on the frequencies and inverse document frequencies corresponding to each of the keywords, use the term frequency-inverse document frequency algorithm to search for research record data that matches the keywords between the target paper and the research record data as the target record data in each of the research record data.
[0058] The term frequency-inverse document frequency algorithm (TF-IDF) is an algorithm commonly used in information retrieval and text mining. Its core idea is to calculate the importance of a word in a document in order to rank and recommend documents in applications such as search engines.
[0059] In order to further improve the effectiveness and reliability of obtaining the time dimension score, in a paper review method based on multi-dimensional data fusion provided in an embodiment of the present application, see Figure 2 , step 100 in the paper review method based on multi-dimensional data fusion further specifically includes the following content: Step 140: Extract the corresponding submission time from the software system to obtain the submission time period corresponding to all the target record data; determine the time dimension score corresponding to the target paper for representing the time distribution of the paper research according to the submission time period and a preset empirical threshold.
[0060] Specifically, the time dimension score S can be calculated according to formula (1): Formula (1) Where refers to the scoring function, and the value of the scoring function after solution is used as the time dimension score S; is the submission time period corresponding to all the target record data, that is, the time period between the submission time when the target record data is first detected in the software system and the submission time when the target record data is last detected; is the empirical threshold, and here the present application takes days, days. In formula (1), when (that is, completed instantly, an extreme case of temporary piecing together), points; when is the case, points; when is the case, 10 points.
[0061] Step 150: Based on the target record data and the paper text data corresponding to the target paper, obtain the spatial variable corresponding to the target paper for representing the spatial distribution of the paper research.
[0062] And, step 160: According to the target record data and the paper text data corresponding to the target paper, obtain the text correlation coefficient corresponding to the target paper for representing the similarity between the paper text data and the target record data.
[0063] To further improve the effectiveness and reliability of obtaining spatial variables, in a paper review method based on multi-dimensional data fusion provided in an embodiment of the present application, refer to Figure 3 , step 150 in the paper review method based on multi-dimensional data fusion specifically includes the following content: Step 151: Obtain the check-in information corresponding to each of the target record data submitted by the user who is the author of the target paper in the software system, and respectively determine the geographical location information corresponding to each of the check-in information according to the geographic information system.
[0064] Step 152: Perform periodic encoding on each of the geographical location information to obtain a spatial variable corresponding to the target paper for representing the spatial distribution of the paper research.
[0065] The core purpose of periodic encoding is to convert periodic features into a form recognizable by the model through mathematical methods (such as sine or cosine functions) to capture repetitive patterns and multi-scale location information.
[0066] To further improve the effectiveness and reliability of obtaining the text correlation coefficient, in a paper review method based on multi-dimensional data fusion provided in an embodiment of the present application, refer to Figure 3 , step 160 in the paper review method based on multi-dimensional data fusion specifically includes the following content: Step 161: Generate summary text data corresponding to the target record data based on a large language model.
[0067] Step 162: Input the summary text data and the paper text data into the BERT model respectively, so that the BERT model outputs the summary feature vector corresponding to the summary text data and the paper feature vector corresponding to the paper text data respectively.
[0068] The BERT model adopts the bidirectional self-attention mechanism of the Transformer encoder to simultaneously capture left and right context information in all layers and achieve richer semantic representations.
[0069] Step 163: Calculate the similarity between the summary feature vector and the paper feature vector using cosine similarity as the text correlation coefficient corresponding to the target paper for representing the similarity between the paper text data and the target record data.
[0070] In order to further improve the effectiveness and reliability of obtaining content quality feature vectors, in a paper review method based on multidimensional data fusion provided in an embodiment of the present application, see Figure 3 , step 200 in the paper review method based on multidimensional data fusion specifically includes the following contents: Step 210: preprocessing the paper text data; wherein the preprocessing includes: removing noise, sentence segmentation and word segmentation.
[0071] Step 220: Based on the corresponding word segmenter of natural language processing technology, obtain the word embedding vector sequence corresponding to the preprocessed paper text data.
[0072] Step 230: Linearly map the hidden feature vectors corresponding to the word embedding vector sequence to obtain a content quality feature vector corresponding to the target paper for representing the content quality of the paper.
[0073] Specifically, the SciBERT model can be used for deep semantic encoding. The input paper text is standardized, including noise removal (such as extra spaces, punctuation cleaning, etc.), sentence segmentation, and word segmentation. Then, the text is converted into a subword sequence using the SciBERT built-in word segmenter to ensure that the input format meets the model requirements. After passing through the SciBERT word segmenter, the vocabulary sequence is obtained and then converted into the corresponding word embedding vector sequence.
[0074] The SciBERT model is a pre-trained model based on the BERT model, specifically for scientific text processing. It improves the performance of natural language processing tasks in the scientific field by pre-training on a large amount of scientific literature.
[0075] In order to further improve the effectiveness and reliability of obtaining the text quantity index, in a paper review method based on multidimensional data fusion provided in an embodiment of the present application, see Figure 3 , step 200 in the paper review method based on multidimensional data fusion further specifically includes the following contents: Step 240: Obtain the data volume, number of experiments and number of cited literature corresponding to the paper text data.
[0076] Step 250: Standardize the data volume, number of experiments and number of cited documents respectively to obtain a first standard value corresponding to the data volume, a second standard value corresponding to the number of experiments and a third standard value corresponding to the number of cited documents.
[0077] Specifically, the first standard value corresponding to the data volume Z The calculation formula is shown in formula (2): Formula (2) Among them, refers to the mean value corresponding to the data volume, refers to the standard deviation corresponding to the data volume; The second standard value corresponding to the number of experiments E The calculation formula is as shown in formula (2): Formula (3) Among them, refers to the mean value corresponding to the number of experiments, refers to the standard deviation corresponding to the number of experiments; The number of cited references The corresponding third standard value The calculation formula is as shown in formula (4): Formula (4) Among them, refers to the mean value corresponding to the number of cited references, refers to the standard deviation corresponding to the number of cited references.
[0078] Step 260: Calculate the original quantity index corresponding to the thesis text data according to the first standard value, the second standard value, and the second standard value and their respective weights.
[0079] Specifically, the original quantity index corresponding to the thesis text data The calculation formula is as shown in formula (5): Formula (5) Among them, the weight of the data volume is , the weight of the number of experiments is , the weight of the number of cited references is , and .
[0080] Step 270: Perform a linear transformation on the original quantity index to convert it to the [0, 10] interval, and obtain the text quantity index for the target thesis to represent the proportion of quantity dimension information in the thesis.
[0081] The calculation formula of the text quantity index N is as shown in formula (6): Formula (6) Among them, the original quantity index The value range of is ; The linear transformation formula converts the score to the [0, 10] interval to obtain the value of N.
[0082] To further improve the effectiveness and reliability of paper review based on multi-dimensional data fusion, in a paper review method based on multi-dimensional data fusion provided in an embodiment of the present application, refer to Figure 4 The multi-dimensional data fusion paper review model based on the interactive attention mechanism in the paper review method based on multi-dimensional data fusion specifically includes the following contents: A scalar embedding layer, configured to perform scalar embedding on the time dimension score S, the text correlation coefficient k, and the text quantity index N corresponding to the target paper, so as to obtain a time feature vector S1 corresponding to the time dimension score, a text correlation feature vector k1 corresponding to the text correlation coefficient, and a text quantity feature vector N1 corresponding to the text quantity index; A vector projection layer, configured to perform vector projection on the spatial variable D and the content quality feature vector Q corresponding to the target paper, so as to obtain a projected spatial variable D1 corresponding to the spatial variable and a projected content quality feature vector Q1 corresponding to the content quality feature vector; An input sequence construction layer, configured to construct a corresponding input sequence according to the time feature vector S1, the text correlation feature vector k1, the text quantity feature vector N1, the projected spatial variable D1, and the projected content quality feature vector Q1; A spatial and content cross-attention layer, configured to use the projected content quality feature vector Q1 in the input sequence as a query vector, and use the projected spatial variable D1 as a key-value pair vector to perform cross-attention calculation, so as to obtain a corresponding spatial and content fusion feature vector D_Q; A quantity and quality cross-attention layer, configured to splice the projected content quality feature vector Q1 and the spatial and content fusion feature vector D_Q in the input sequence to obtain a corresponding first spliced vector, that is: Q + D_Q; use the first spliced vector Q + D_Q as a query vector, and use the text quantity feature vector N1 in the input sequence as a key-value pair vector to perform cross-attention calculation, so as to obtain a corresponding quantity and quality fusion feature vector N_Q; A vector splicing layer, configured to splice the projected content quality feature vector Q1, the projected spatial variable D1, and the quantity and quality fusion feature vector N_Q in the input sequence to obtain a corresponding second spliced vector; A dynamic fusion layer, configured to perform dynamic fusion on the second spliced vector by using the text correlation feature vector k1 in the input sequence as a modulation factor; A three-dimensional attention routing layer, configured to obtain a target output representation corresponding to the target paper according to the self-interaction result data among the time feature vector S1, the text correlation feature vector k1, and the text quantity feature vector N1, the self-interaction result data between the projected content quality feature vector Q1 and the projected spatial variable D1, and the second concatenated vector after dynamic fusion; A linear layer, configured to map the target output representation to a paper review score corresponding to the target paper.
[0083] From a software perspective, the present application further provides a paper review device based on multi-dimensional data fusion for executing all or part of the content in the paper review method based on multi-dimensional data fusion. Refer to Figure 5 , the paper review device based on multi-dimensional data fusion specifically includes the following content: A first multi-dimensional data acquisition module 10, configured to extract research record data matching the current target paper from a software system for storing research record data and clock-in information periodically submitted by users as target record data, and respectively obtain a time dimension score S representing the time distribution of the target paper research, a spatial variable D representing the spatial distribution of the target paper research, and a text correlation coefficient k representing the similarity between the paper text data and the target record data according to the target record data and the paper text data corresponding to the target paper.
[0084] And a second multi-dimensional data acquisition module 20, configured to determine a content quality feature vector Q representing the content quality of the target paper according to the paper text data based on natural language processing technology, and determine a text quantity index N representing the proportion of the quantity dimension information in the target paper according to the paper text data.
[0085] A paper review scoring module 30, configured to input the time dimension score S, the text correlation coefficient k, the text quantity index N, the spatial variable D, and the content quality feature vector Q corresponding to the target paper into a multi-dimensional data fusion paper review model based on an interactive attention mechanism, so that the paper review model outputs a paper review score corresponding to the target paper.
[0086] The embodiment of the paper review device based on multi-dimensional data fusion provided by the present application can specifically be used to execute the processing flow of the embodiment of the paper review method based on multi-dimensional data fusion in the above embodiment, and its functions will not be described in detail here. For details, reference can be made to the detailed description of the embodiment of the paper review method based on multi-dimensional data fusion above.
[0087] The part of the paper review device based on multi-dimensional data fusion for conducting paper review based on multi-dimensional data fusion can be completed in a server or a client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor for the specific processing of paper review based on multi-dimensional data fusion.
[0088] The above-mentioned client device may have a communication module (i.e., communication unit), which can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side. In other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform having a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.
[0089] Any suitable network protocol can be used for communication between the above-mentioned server and the client device, including network protocols not yet developed on the filing date of this application. The network protocol may, for example, include TCP / IP protocol, UDP / IP protocol, HTTP protocol, HTTPS protocol, etc. Of course, the network protocol may also, for example, include RPC protocol (Remote Procedure Call Protocol) and REST protocol (Representational State Transfer) used on top of the above protocols, etc.
[0090] As can be seen from the above description, the paper review device based on multi-dimensional data fusion provided by the embodiments of the present application extracts the research record data periodically submitted by the user that matches the current target paper as the target record data from a software system for storing the research record data and check-in information periodically submitted by the user, and respectively obtains the time dimension score, spatial variable, and text correlation coefficient corresponding to the target paper according to the target record data and the paper text data corresponding to the target paper, which can effectively trace the user's daily paper research time, research space distribution, and similarity, and can effectively improve the recognition accuracy of whether the paper content is pieced together temporarily; and, by using the time dimension score S, the text correlation coefficient k, the text quantity index N, the spatial variable D, and the content quality feature vector Q corresponding to the target paper as multi-dimensional data for automatic scoring of the paper review score, the rationality and effectiveness of the paper review score can be improved; that is to say, the present application can improve the accuracy of the temporary piecing judgment of the paper review without using blockchain technology, and thus can effectively reduce the resource occupancy rate and storage cost of the device for performing the paper review and lower the technical threshold; and compared with the method of using an AI paper quality portrait, it can effectively improve the rationality and effectiveness of the paper review score and reduce the misjudgment rate.
[0091] To further illustrate the above embodiments, the present application also provides a specific application example of a paper review method based on multi-dimensional data fusion. In this application example, taking a software system as the postgraduate monthly report system (which can be abbreviated as the monthly report system or the campus management system, etc.) as an example, correspondingly, the research record data can be expressed as a monthly report. Based on this, the application example of the present application specifically includes the following contents: (1) Time data tracing based on the postgraduate monthly report system, scoring function based on the time of the postgraduate monthly report system Obtain the time dimension score S: Using the time dimension data such as the postgraduate monthly report submission time and content update frequency recorded by the postgraduate monthly report system, and applying the term frequency-inverse document frequency (TF-IDF) method, extract the keywords in the monthly report and the paper. First, calculate the frequency of the keyword appearing in the text : Among them, represents the keyword in the document The number of times it appears; represents the index variable, which is used to traverse all different words appearing in the document j; represents the number of times all different words appear; represents the keyword in the document The frequency of appearance in.
[0092] Then calculate the generality of the keywords, i.e., the inverse document frequency : Among them, represents the number of all documents, represents the number of documents containing the keyword . If the number of documents containing the i-th keyword is smaller, is larger, it indicates that the term has good category discrimination ability.
[0093] A high term frequency within a specific document and a low document frequency of the term in the entire document collection can produce a high-weight TF-IDF: Among them, represents the TF-IDF value of the keyword in the document ; represents the frequency of the keyword appearing in the document , represents the inverse document frequency.
[0094] Using the TF-IDF method, the keywords in the monthly reports and papers can be extracted, the relevant keywords can be compared, the relevant monthly reports with a similarity of more than 80% to the paper keywords can be extracted, and the submission time of the monthly reports can be retrieved from the background.
[0095] According to whether the content is temporarily pieced together or accumulated over a long time, based on the scoring function of the time of the postgraduate monthly report system calculate the time dimension score S, as shown in the aforementioned formula (1).
[0096] According to the keyword matching results, analyze the time distribution of the research content related to the paper. Judge whether the research shows the characteristics of temporary piecing together with a concentrated outbreak in a short time or the result of continuous accumulation and gradual improvement over a long time. And assign a corresponding value of 0 - 10 points to the paper.
[0097] (2) Integrate the verification of spatial information of institutions such as laboratories to construct the spatial variable D: With the help of the Geographic Information System (GIS) and the school's internal management system, combined with spatial dimension information such as the site usage records and equipment borrowing registrations of the graduate school and research rooms, verify whether the thesis research work is carried out in the relevant places where the thesis submitter is located. At the same time, for cases involving cooperation, check cooperation agreements, communication records, etc. to confirm the authenticity and participation of the cooperation, so as to prove the close connection between the research work and the thesis submitter. The Geographic Information System is a computer system used for collecting, storing, managing, analyzing, and displaying geospatial data.
[0098] The monthly report system for postgraduate students requires students to check in regularly, and the IP geographical location information of each check-in can be obtained. Taking the longitude and latitude information (x, y) collected by the GIS system as input, this application can use a periodic coding to represent: There is a continuous spatial variable , in order to capture its periodic information, this application pre-defines a set of frequency parameters , where each , , where is the initial frequency.
[0099] Then the periodic coding vector of can be defined as: For the two-dimensional spatial variable , the periodic coding can be performed on x and y respectively, and then the two coding vectors are concatenated: This formal expression uses sine and cosine functions to map the continuous spatial variable to a high-dimensional periodic feature space, which can retain the periodic structure of the original spatial data and is convenient to be used as part of the feature vector in the subsequent automated scoring model.
[0100] (3) Use text analysis technology to evaluate the content quality and construct a content quality feature vector : Use natural language processing (NLP) technology to deeply analyze the thesis text. Natural Language Processing is an important branch in the fields of artificial intelligence and computer science, mainly studying how to make computers understand and process human natural language.
[0101] The SciBERT model is used for deep semantic encoding. The input paper text is standardized, including noise removal (such as extra spaces, punctuation cleaning, etc.), sentence segmentation, and word segmentation. Then the text is converted into a subword sequence using the SciBERT built-in word segmenter to ensure that the input format meets the model requirements. Let the preprocessed text be T, and the vocabulary sequence is obtained after passing through the SciBERT word segmenter, and then converted into the corresponding word embedding vector sequence. After processing by the SciBERT model, the hidden layer representation corresponding to the output [CLS] tag is recorded as: in, represents the feature vector representing the classification head; Represents a real vector with a dimension of 768, that is, this vector contains 768 real elements Then, through a linear mapping: Among them, W represents the weight matrix, which is used to transfer and transform information between different layers of the SciBERT model; It means that W is a real matrix with shape of 256 rows and 768 columns.
[0102] The 768-dimensional vector is mapped to 256 dimensions to obtain the feature vector Q representing the content quality.
[0103] (4) Construct a comprehensive quality and quantity evaluation model and construct a text quantity index : Based on the results of text analysis, a comprehensive evaluation model is constructed in combination with quantitative dimension information such as the amount of data involved in the paper, the number of experiments, and the number of cited literature. By setting reasonable weights, the quality and quantity of the paper are quantitatively evaluated to comprehensively measure the performance of the paper at the data foundation level.
[0104] Assume that the amount of data involved in the paper is , the number of experiments is The number of cited papers is First, each indicator is standardized to eliminate the dimensional effects between different indicators, as shown in the above formulas (2) to (4).
[0105] in, , , are the means of the amount of data, number of experiments, and number of cited literature, respectively. , , are their standard deviations respectively. Assign weights to each indicator, and the weights are determined according to the characteristics and importance of the research field. If the research field is relatively new and there are few references, then the weight of the data volume should be appropriately set to , the weight of the number of experiments is , the weight of the number of cited references is , and .
[0106] The original number index of the paper can be obtained through the aforementioned formula (5).
[0107] Suppose the original quality score calculated by the above formula (5) , where and are set to -5 and 5 respectively according to experience, and the linear transformation formula is used to convert the score to the range of [0, 10], that is, the text quantity index is expressed as: (5) Conduct correlation analysis on the comparison monthly report content to construct the text correlation coefficient k: To quantify the correlation between the paper and its corresponding monthly report content, this paper proposes an automated method based on large model abstract generation and BERT series model feature extraction. The specific steps are as follows: To reduce redundant information and capture the core content of the monthly report, first use a large-scale pre-trained language model (such as GPT-4 or T5) to generate an abstract of the original monthly report text to obtain the abstract text
[0108] The abstract and the full text of the paper are respectively input into the pre-trained BERT series model, and its deep semantic encoding ability is used to extract the corresponding feature representations. Suppose the feature vectors obtained through the BERT model are respectively: where is the feature vector of the abstract text ; The feature vector of the full text T of the paper.
[0109] The cosine similarity is used to calculate the similarity between the two feature vectors to measure the correlation degree between the monthly report abstract and the paper content. The formula is: where is the feature vector of the monthly report abstract, is the feature vector of the paper text, That is, the text correlation coefficient, whose value range is [0, 1]. The closer the value is to 1, the higher the content correlation between the two.
[0110] (6) Draw a final conclusion by synthesizing the multi-dimensional review results: Integrating the information in multiple aspects such as the time dimension, space dimension, and data basis obtained in the previous five steps, this application can construct an end-to-end model, that is, a multi-dimensional data fusion paper review model based on the interactive attention mechanism. There is a large amount of data in the system of this application where experts score papers. This application uses this data as supervision to train the multi-dimensional data fusion paper review model based on the interactive attention mechanism of this application. The specific results of the multi-dimensional data fusion paper review model based on the interactive attention mechanism are as Figure 6 shown.
[0111] The input of the multi-dimensional data fusion paper review model based on the interactive attention mechanism includes: 1. Time dimension score S: a scalar (0 - 10 points); 2. Space variable D: a high-dimensional vector after periodic encoding (the dimension depends on the number of frequency parameters, and the dimension is 256 when n = 64); 3. Content quality feature vector Q: a 256-dimensional semantic vector; 4. Text quantity index N: a scalar (0 - 10 points); 5. Text correlation coefficient k: a scalar (0 - 1).
[0112] In this model, first, the five input features of the time dimension score S, space variable D, content quality Q, quantity index N, and text correlation k are respectively subjected to scalar embedding or vector projection, and then they are concatenated with the [CLS] token to form a unified input sequence; then, the model uses Q1 as the query vector (Query) and D1 as the key / value pair vector (Key / Value) through "spatial and content cross-attention" to learn the dependence relationship between content quality and space variables, and obtains the fused D_Q; subsequently, "quantity and quality cross-attention" is used to take Q + D_Q as the Query and N as the Key / Value to further introduce the quantity index and obtain N_Q; then the original D1, Q1, and N_Q are concatenated together, and dynamic fusion is performed through the text correlation k as a modulation factor to balance the importance of spatio-temporal features and quantity features; finally, the model uses "three-dimensional attention routing" to simultaneously consider the interactions between scalar features (S1, N1, k1), vector features (D1, Q1), and cross-scalar and vector interactions, and comprehensively obtains the target output representation that can be used for downstream tasks. Finally, this application passes the [CLS] token through a linear layer and maps it to a paper review score. The [CLS] token is used to represent the semantic information of the entire text.
[0113] In addition, to further improve the application effectiveness and reliability of the above application examples, the usage process of the postgraduate monthly report system can be referred to Figure 7 , where the process for students to submit monthly or annual report information is as Figure 8 shown, and the process for the management end to review the monthly or annual report information submitted by students is as Figure 9 shown.
[0114] Based on this, the application example of this application relates to a method for multi-dimensional data fusion paper submission review based on NLP, which is a method for macroscopically controlling the quality of submitted papers by combining AI algorithms. It relates to the fields of postgraduate training and submission management, and adopts the following steps: (1) Generate a scoring function S(T) based on the time data traceability of the postgraduate monthly report system; (2) Integrate the verification of spatial information of institutions such as laboratories to construct a spatial variable D; (3) Use text analysis technology to evaluate the content quality and construct a content quality evaluation index Q; (4) Construct a comprehensive quality and quantity evaluation model and construct a text quantity index N; (5) Conduct correlation analysis by comparing the monthly report content and construct a text correlation coefficient k; (6) Introduce an expert review and feedback mechanism and construct an expert evaluation index P; (7) Draw a final conclusion based on the multi-dimensional review results. This application proposes a method for postgraduate thesis quality management that combines process management and target management, solves the problem that traditional thesis review methods focus more on the finished thesis and pay insufficient attention to the research process, uses natural language processing technology and text similarity algorithms to connect the thesis with the author's daily research accumulation, ensures that the research is based on a solid process, and is conducive to discovering truly valuable and high-quality theses.
[0115] That is to say, the effects of the above application examples of this application are as follows: (1) Currently, the commonly used thesis review methods mainly focus on the finished thesis and pay insufficient attention to the research process. The new method can judge whether the research is pieced together temporarily or accumulated over a long time through the time data traceability of the journal system; by verifying the institutional spatial information, it can verify the research venue and cooperation situation. This all-round and multi-dimensional monitoring of the research process makes up for the defect that traditional methods only focus on the results, guarantees the quality of the thesis from the source, ensures that the research is based on a solid process rather than a hasty patchwork later.
[0116] (2) The evaluation of the content of papers by common review methods mostly relies on the subjective judgment of experts and lacks precise quantification. The new method uses text analysis technology to evaluate the basic quality, and constructs a comprehensive evaluation model by combining quantitative dimension information such as the amount of data and the number of experiments. It takes into account both the quality dimensions of papers, such as vocabulary and logic, and the quantitative dimensions, and realizes quantitative evaluation through reasonable weight setting. In contrast, this method is more scientific and objective, reduces the evaluation deviation caused by subjective factors, and can more accurately measure the performance of papers at the data basis level.
[0117] (3) Traditional review methods rarely link papers with the author's daily research accumulation. The new method conducts correlation analysis by comparing the daily accumulation content, and uses the text similarity algorithm to judge whether the paper integrates the usual research ideas and results, ensuring the accumulation and coherence of the paper content. At the same time, expert review is introduced to evaluate from multiple perspectives such as academic value, and conclusions are drawn by synthesizing the review results of multiple dimensions. This makes the judgment of papers more comprehensive and in-depth, not only paying attention to the current achievements, but also attaching importance to the precipitation of the long-term research process, which is conducive to excavating truly valuable and high-quality papers.
[0118] The embodiment of the present application also provides an electronic device, which may include a processor, a memory, a receiver and a transmitter. The processor is used to execute the paper review method based on multi-dimensional data fusion mentioned in the above embodiment. The processor and the memory may be connected through a bus or other means. Taking the connection through the bus as an example. The receiver can be connected to the processor and the memory in a wired or wireless manner.
[0119] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.
[0120] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs and modules, such as the program instructions / modules corresponding to the paper review method based on multi-dimensional data fusion in the embodiment of the present application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions and modules stored in the memory, that is, to implement the paper review method based on multi-dimensional data fusion in the above method embodiment.
[0121] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0122] The one or more modules are stored in the memory, and when executed by the processor, perform the paper review method based on multidimensional data fusion in the embodiment.
[0123] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit, which may include a receiver and a transmitter. The processor, memory, receiver, and transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to send and receive signals.
[0124] As an implementation method, the functions of the receiver and the transmitter in the present application can be considered to be implemented through a transceiver circuit or a dedicated chip for transceiver, and the processor can be considered to be implemented through a dedicated processing chip, a processing circuit or a general chip.
[0125] As another implementation method, it is possible to use a general-purpose computer to implement the server provided in the embodiment of the present application, that is, to store the program code for implementing the functions of the processor, receiver, and transmitter in a memory, and the general-purpose processor implements the functions of the processor, receiver, and transmitter by executing the code in the memory.
[0126] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the paper review method based on multidimensional data fusion are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.
[0127] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the aforementioned paper review method based on multidimensional data fusion.
[0128] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or a communication link.
[0129] It should be clear that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of this application.
[0130] In this application, features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0131] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, various changes and variations can be made to the embodiments of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A paper review method based on multi-dimensional data fusion, characterized in that Including: From a software system for storing research record data and check-in information periodically submitted by users, extracting research record data matching the current target paper as target record data, and respectively obtaining, according to the target record data and the paper text data corresponding to the target paper, a time dimension score for representing the time distribution of the paper research, a spatial variable for representing the spatial distribution of the paper research, and a text correlation coefficient for representing the similarity between the paper text data and the target record data; And, based on natural language processing technology, determining, according to the paper text data, a content quality feature vector for representing the content quality of the target paper, and determining, according to the paper text data, a text quantity index for representing the proportion of quantity dimension information in the target paper; Inputting the time dimension score, the text correlation coefficient, the text quantity index, the spatial variable, and the content quality feature vector corresponding to the target paper into a multi-dimensional data fusion paper review model based on an interactive attention mechanism, so that the paper review model outputs a paper review score corresponding to the target paper.
2. The method for reviewing papers based on multi-dimensional data fusion according to claim 1, wherein In the software system for storing research record data and check-in information periodically submitted by users, extracting research record data matching the current target paper as target record data includes: In the software system for storing research record data and check-in information periodically submitted by users, extracting each research record data periodically submitted by the user who is the author of the target paper; Respectively extracting the paper text data corresponding to the target paper and the keywords corresponding to each of the research record data, and calculating the frequency and inverse document frequency of each keyword appearing in the paper text data or the research record data where it is located; According to the frequency and inverse document frequency corresponding to each keyword, using the term frequency-inverse document frequency algorithm to find, in each of the research record data, research record data with keyword matching with the target paper as target record data.
3. The method for reviewing papers based on multi-dimensional data fusion according to claim 1, characterized in that, The step of respectively obtaining, according to the target record data and the paper text data corresponding to the target paper, a time dimension score for representing the time distribution of the paper research, a spatial variable for representing the spatial distribution of the paper research, and a text correlation coefficient for representing the similarity between the paper text data and the target record data includes: Extracting the submission time corresponding to the target record data from the software system to obtain the submission time period corresponding to all the target record data; determining, according to the submission time period and a preset empirical threshold, a time dimension score for representing the time distribution of the paper research corresponding to the target paper; According to the target record data and the paper text data corresponding to the target paper, obtaining a spatial variable for representing the spatial distribution of the paper research corresponding to the target paper; And, according to the target record data and the paper text data corresponding to the target paper, obtain the text correlation coefficient corresponding to the target paper, which is used to represent the similarity between the paper text data and the target record data.
4. The method for reviewing papers based on multi-dimensional data fusion according to claim 3, characterized in that, The obtaining of the spatial variable corresponding to the target paper, which is used to represent the spatial distribution of the paper research according to the target record data and the paper text data corresponding to the target paper, includes: Obtain the check-in information corresponding to each of the target record data submitted by the user who is the author of the target paper in the software system, and respectively determine the geographical location information corresponding to each of the check-in information according to the geographic information system; Perform periodic encoding on each of the geographical location information to obtain the spatial variable corresponding to the target paper, which is used to represent the spatial distribution of the paper research.
5. The method for reviewing papers based on multi-dimensional data fusion according to claim 3, characterized in that The obtaining of the text correlation coefficient corresponding to the target paper, which is used to represent the similarity between the paper text data and the target record data according to the target record data and the paper text data corresponding to the target paper, includes: Generate the abstract text data corresponding to the target record data based on the large language model; Input the abstract text data and the paper text data into the BERT model respectively, so that the BERT model outputs the abstract feature vector corresponding to the abstract text data and the paper feature vector corresponding to the paper text data respectively; Adopt cosine similarity to calculate the similarity between the abstract feature vector and the paper feature vector, and use it as the text correlation coefficient corresponding to the target paper, which is used to represent the similarity between the paper text data and the target record data.
6. The method for reviewing papers based on multi-dimensional data fusion according to claim 1, wherein, The determining of the content quality feature vector corresponding to the target paper, which is used to represent the content quality of the paper based on natural language processing technology, includes: Preprocess the paper text data; wherein, the preprocessing includes: removing noise, clause segmentation and word segmentation; Based on natural language processing technology corresponding to the tokenizer, obtain the word embedding vector sequence corresponding to the preprocessed paper text data; Perform linear mapping on the hidden layer feature vector corresponding to the word embedding vector sequence to obtain the content quality feature vector corresponding to the target paper, which is used to represent the content quality of the paper.
7. The method for paper review based on multi-dimensional data fusion according to claim 1, characterized in that The determining of the text quantity index corresponding to the target paper, which is used to represent the proportion of the quantity dimension information in the paper according to the paper text data, includes: Obtain the data volume, the number of experiments and the number of cited documents corresponding to the paper text data; Perform standardization processing on the data volume, the number of experiments and the number of cited documents respectively to obtain the first standard value corresponding to the data volume, the second standard value corresponding to the number of experiments and the third standard value corresponding to the number of cited documents; Calculate the original quantity index corresponding to the paper text data according to the first standard value, the second standard value and the second standard value and their respective corresponding weights; Perform a linear transformation on the original quantity metric to convert it to the range [0, 10], obtaining the text quantity metric corresponding to the target paper, which is used to represent the proportion of quantity dimension information in the paper.
8. The method for reviewing papers based on multi-dimensional data fusion according to any one of claims 1 to 7, characterized in that The multi-dimensional data fusion paper review model based on the interactive attention mechanism includes: A scalar embedding layer for performing scalar embedding on the time dimension score, the text correlation coefficient, and the text quantity metric corresponding to the target paper to obtain a time feature vector corresponding to the time dimension score, a text correlation feature vector corresponding to the text correlation coefficient, and a text quantity feature vector corresponding to the text quantity metric; A vector projection layer for performing vector projection on the spatial variable and the content quality feature vector corresponding to the target paper to obtain a projected spatial variable corresponding to the spatial variable and a projected content quality feature vector corresponding to the content quality feature vector; An input sequence construction layer for constructing a corresponding input sequence according to the time feature vector, the text correlation feature vector, the text quantity feature vector, the projected spatial variable, and the projected content quality feature vector; A spatial and content cross-attention layer for using the projected content quality feature vector in the input sequence as a query vector and the projected spatial variable as a key-value pair vector for cross-attention calculation to obtain a corresponding spatial and content fusion feature vector; A quantity and quality cross-attention layer for concatenating the projected content quality feature vector in the input sequence with the spatial and content fusion feature vector to obtain a corresponding first concatenated vector; using the first concatenated vector as a query vector and the text quantity feature vector in the input sequence as a key-value pair vector for cross-attention calculation to obtain a corresponding quantity and quality fusion feature vector; A vector concatenation layer for concatenating the projected content quality feature vector, the projected spatial variable, and the quantity and quality fusion feature vector in the input sequence to obtain a corresponding second concatenated vector; A dynamic fusion layer for dynamically fusing the second concatenated vector with the text correlation feature vector in the input sequence as a modulation factor; A three-dimensional attention routing layer for obtaining the target output representation corresponding to the target paper according to the self-interaction result data between the time feature vector, the text correlation feature vector, and the text quantity feature vector, the self-interaction result data between the projected content quality feature vector and the projected spatial variable, and the second concatenated vector after dynamic fusion; A linear layer for mapping the target output representation to the paper review score corresponding to the target paper.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the paper review method based on multi-dimensional data fusion according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the paper review method based on multi-dimensional data fusion according to any one of claims 1 to 8.
Citation Information
Patent Citations
Text processing method and device, electronic equipment and computer readable storage medium
CN113011126A
Paper reviewer determination method and system
CN114154478A
Method and system for submission, evaluation, publication of research article and citation index calculation in on-line
KR1020170117781A
Automated essay scoring
US20040175687A1
Information technology networked entity monitoring with automatic reliability scoring
US20190095478A1
Cited By
Paper quality intelligent evaluation system based on multi-model fusion
CN120430291A
Factory paper detection method, system and equipment based on migration characteristics
CN121189303A