Cross-domain literature value evaluation method and system based on semantic element modeling
By constructing a hierarchical cross-domain literature database and training a value transfer weight mapping model, the problems of high sample dependence and poor cross-domain adaptability in academic literature retrieval are solved, and efficient and accurate screening and recommendation of cross-domain literature is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies for academic literature retrieval suffer from high sample dependence, poor cross-domain adaptability, and shallow semantic understanding, making it difficult to accurately obtain literature in the same field within a specific subfield and causing the loss of high-value cross-domain resources.
By constructing a hierarchical cross-domain literature database, performing domain-level deep annotation and hierarchical index storage, extracting multi-dimensional semantic elements using semantic element extraction templates, and training a value transfer weight mapping model, cross-domain literature value assessment and screening can be carried out.
It enables precise screening and efficient recommendation of high-value academic literature across disciplines, improving the efficiency and quality of literature resource utilization.
Smart Images

Figure CN121542446B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of academic literature mining, and in particular to cross-domain literature value assessment methods and systems based on semantic element modeling. Background Technology
[0002] Academic literature retrieval and value screening are crucial for the efficient conduct of interdisciplinary research activities, directly impacting researchers' efficiency and the quality of their findings. Especially in specific research areas, accurate acquisition of literature resources is a core prerequisite for overcoming research bottlenecks. Currently, the main methods for addressing this problem are behavior-driven recommendation based on collaborative filtering, content-driven recommendation based on keywords or surface text features, and traditional citation tracing. These methods either rely on interactive data and sufficient samples within the same field or only match literature through shallow features. They are unable to address the scarcity of samples within specific fields and lack the ability to mine cross-domain semantic connections. This results in a double dilemma: difficulty finding sufficient relevant literature within the same field and easy loss of semantically relevant high-value resources in other fields, creating a situation where neither relevant literature within the same field nor cross-domain resources can be found.
[0003] At present, the relevant technologies for literature value assessment have technical problems such as high sample dependence, poor cross-domain adaptability, and shallow semantic understanding. Summary of the Invention
[0004] This application provides a method and system for cross-domain literature value assessment based on semantic element modeling. It constructs a hierarchical cross-domain literature database by performing domain-depth annotation and hierarchical indexing of multi-domain literature. Based on the domain depth of the current literature, it selects cross-domain literature of the same level from the database to form a dedicated literature database. Using semantic element extraction templates, it extracts multi-dimensional semantic elements from both the current literature and the cross-domain literature of the same level database. A value transfer weight mapping model is trained to obtain the corresponding value transfer weights for each domain. The two types of semantic elements are aligned according to these weights, and the identification of literature whose first-class value exceeds a preset threshold is calculated and selected. This completes the assessment and selection of high-value cross-domain literature. These technical means solve the technical problems of high sample dependence, poor cross-domain adaptability, and shallow semantic understanding in existing literature value assessment methods, achieving the technical effect of accurate selection and efficient recommendation of high-value cross-domain academic literature.
[0005] This application provides a method for cross-domain literature value assessment based on semantic element modeling, comprising: constructing a hierarchical cross-domain literature database, wherein the hierarchical cross-domain literature database is indexed and stored hierarchically according to the domain depth of the documents in the database by performing domain depth annotation on the documents in the database; obtaining a cross-domain peer-level literature database from the hierarchical cross-domain literature database according to the domain depth of the current document; obtaining multidimensional semantic elements of the current document and the multidimensional semantic element database in the cross-domain peer-level literature database through a semantic element extraction template; training a value transfer weight mapping model, wherein the value transfer weight mapping model is used to perform value transfer weight mapping on cross-domain documents to obtain the value transfer weight under each domain; aligning the multidimensional semantic elements and the multidimensional semantic element database with transfer weights according to the value transfer weight under each domain, and outputting the identified documents whose first document value under each domain is greater than a preset document value threshold.
[0006] In a possible implementation, the documents in the database are hierarchically indexed and stored according to the labeled domain depth, and the following processing is performed: keyword vectors and key sentence vectors are extracted from the documents in the database to obtain a set of keyword and sentence vectors; multi-level domain depth labels and corresponding multi-level word and sentence vector sets are defined, wherein the multi-level domain depth labels include basic layer depth, method layer depth, application layer depth, and cross layer depth; cosine similarity is introduced to calculate the cosine similarity between the multi-level word and sentence vector sets and the keyword and sentence vector sets, and the domain depth label of each document in its respective domain is output; the documents in the database are hierarchically indexed and stored according to the domain depth label of each document in its respective domain.
[0007] In a possible implementation, the multidimensional semantic elements and the multidimensional semantic element library are aligned according to the value transfer weights under each domain. The following processing is also performed: a semantic vector mapping space is configured, and the multidimensional semantic elements and the multidimensional semantic element library are respectively input into the semantic vector mapping space to obtain multidimensional semantic element vectors and a multidimensional semantic element vector library; the multidimensional semantic element vectors and the multidimensional semantic element vector library are aligned, and document value is evaluated according to the value transfer weights under each domain, outputting a first document value evaluation result; the first document value evaluation result is used to filter identified documents whose first document value under each domain is greater than a preset document value threshold.
[0008] In a possible implementation, a value transfer weight mapping model is trained by performing the following processes: constructing multi-domain comparison literature samples and labeling the cross-domain citation counts of the multi-domain comparison literature samples; extracting multi-dimensional semantic elements from the multi-domain comparison literature samples using a semantic element extraction template to obtain multi-dimensional semantic element comparison samples; initializing value transfer weights and calculating the loss data of the multi-dimensional semantic element comparison samples based on the initial value transfer weights; performing gradient optimization on the initial value transfer weights based on the loss data until the value transfer weights mapped to each domain are obtained, thereby generating a value transfer weight mapping model.
[0009] In a possible implementation, the following processing is performed: calculating the loss data of the multidimensional semantic element comparison samples based on the initial value transfer weights; wherein the loss data includes the cross-domain alignment loss, annotation prediction loss, and cross-domain transferability regularization loss of the semantic elements.
[0010] In a possible implementation, after obtaining the multidimensional semantic elements of the current document and the multidimensional semantic element library in the cross-domain peer-level document library through the semantic element extraction template, the following processing is performed: generating multidimensional extended sentences composed of the multidimensional semantic elements of the current document; generating a multidimensional extended sentence library based on the multidimensional semantic element library; aligning the multidimensional extended sentences and the multidimensional extended sentence library with migration weights according to the value migration weights under each domain, and outputting the identified documents whose second document value is greater than a preset document value threshold under each domain.
[0011] In a possible implementation, the following processing is performed: the first document value is the document value obtained based on multidimensional semantic elements, and the second document value is the document value obtained based on multidimensional extended sentences.
[0012] In a possible implementation, a cross-domain peer literature library is obtained from the hierarchical cross-domain literature library based on the domain depth of the current literature, and the following processing is also performed: when the number of identified documents output based on the cross-domain peer literature library is less than the number required by the user, a cross-domain neighbor literature library is obtained from the cross-domain literature library; and the identified documents output by the cross-domain neighbor literature library are output as an increment.
[0013] In a possible implementation, the first document value in each domain is output, and the following processing is performed: the multidimensional semantic elements and the multidimensional semantic element library are aligned by migration weights, and the cosine similarity of the semantic elements is calculated; the cosine similarity of the semantic elements is weighted according to the value migration weight in each domain to obtain the first document value of each document in each domain.
[0014] This application also provides a cross-domain literature value assessment system based on semantic element modeling, comprising: a hierarchical cross-domain literature database construction module for constructing a hierarchical cross-domain literature database, wherein the hierarchical cross-domain literature database is indexed and stored hierarchically according to the domain depth of the documents in the database by performing domain depth annotation on the documents in the database; a cross-domain peer-level literature database acquisition module for acquiring a cross-domain peer-level literature database from the hierarchical cross-domain literature database according to the domain depth of the current document; a multi-dimensional semantic element acquisition module for acquiring multi-dimensional semantic elements of the current document and the multi-dimensional semantic element database in the cross-domain peer-level literature database through a semantic element extraction template; a value transfer weight mapping model training module for training a value transfer weight mapping model, wherein the value transfer weight mapping model is used to perform value transfer weight mapping on cross-domain documents to obtain the value transfer weight under each domain; and an identification document output module for aligning the multi-dimensional semantic elements and the multi-dimensional semantic element database according to the value transfer weight under each domain, and outputting identification documents in each domain whose first document value is greater than a preset document value threshold.
[0015] This application proposes a method and system for cross-domain literature value assessment based on semantic element modeling. First, a hierarchical cross-domain literature database is constructed. This database is indexed and stored hierarchically according to the domain depth of the documents in the database. Next, a cross-domain peer-level literature database is retrieved from the hierarchical cross-domain literature database based on the domain depth of the current document. Then, multidimensional semantic elements of the current document and the multidimensional semantic element database from the cross-domain peer-level literature database are obtained through a semantic element extraction template. A value transfer weight mapping model is then trained to perform value transfer weight mapping on cross-domain documents, obtaining the value transfer weight for each domain. Finally, the multidimensional semantic elements and the multidimensional semantic element database are aligned according to the value transfer weight for each domain, outputting identified documents whose first document value in each domain is greater than a preset document value threshold. Through the above process, the method and system proposed in this application achieve the technical effect of accurate screening and efficient recommendation of high-value cross-domain academic literature. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings of the embodiments of the present invention will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0017] Figure 1A flowchart illustrating the cross-domain document value assessment method based on semantic element modeling provided in this application embodiment.
[0018] Figure 2 A schematic diagram of the structure of a cross-domain document value assessment system based on semantic element modeling provided in this application embodiment.
[0019] Figure labeling: Module 10 for constructing a hierarchical cross-domain literature database, Module 20 for acquiring a cross-domain peer-level literature database, Module 30 for acquiring multi-dimensional semantic elements, Module 40 for training a value transfer weight mapping model, and Module 50 for identifying literature output. Detailed Implementation
[0020] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0021] This application provides a cross-domain document value assessment method based on semantic element modeling, such as... Figure 1 As shown, the method includes:
[0022] Step S100: Construct a hierarchical cross-domain literature database. The hierarchical cross-domain literature database performs domain depth annotation on the literature in the database and stores the literature in the database hierarchically according to the annotation domain depth.
[0023] Specifically, this involves collecting academic literature from multiple fields, including PDF and text formats. Using literature parsing tools such as Apache PDFBox and MinerU, unstructured documents are converted into structured text, extracting core content such as titles, abstracts, keywords, and main text sections. Domain-level annotation rules are designed based on a domain knowledge system. Annotation tools, such as LabelStudio, combined with large language models like Qwen-7B, are used to perform domain-level annotation on the literature. The annotation process includes characteristics such as the research topic, methodology, and cross-domain relevance of the references. A hierarchical indexing architecture, such as Elasticsearch's multi-level index, is adopted to categorize and store the literature according to the depth of the annotated domain, enabling rapid location of literature sets at specific depths later.
[0024] For example, literature in the field of computer vision that only introduces the basic concepts of convolutional neural networks is labeled as the basic layer depth, while literature on object detection methods based on convolutional neural networks is labeled as the method layer depth, and stored in the corresponding index partitions.
[0025] In one possible implementation, the documents in the database are hierarchically indexed and stored according to the labeled domain depth. Step S100 further includes step S110, which extracts keyword vectors and key sentence vectors from the documents in the database to obtain a set of keyword and sentence vectors. Specifically, keyword extraction uses the TF-IDF algorithm combined with a domain dictionary, such as the ACM Computing Classification System dictionary in the computer science field, to select words with high frequency and strong domain relevance in the documents, such as random forest, semantic vectorization, and Transformer. Key sentence extraction identifies topic sentences, conclusion sentences, and method description sentences in the documents, and uses the TextRank algorithm to calculate sentence importance scores, selecting the top 20% of sentences as key sentences. A unified embedding model, such as Qwen3-Embedding, is used to vectorize the extracted keywords and key sentences, converting the text information into fixed-dimensional real-valued vectors. Finally, the keyword vectors and key sentence vectors of all documents are integrated to form a set of keyword and sentence vectors.
[0026] For example, in the paper "A Text Classification Method Based on Random Forest", keywords such as "random forest, text classification, feature engineering" and key sentences such as "This paper proposes a random forest text classification model that integrates attention mechanism" are extracted, encoded, and stored in the set.
[0027] Step S120: Define multi-level domain depth labels and corresponding multi-level word and phrase vector sets. The multi-level domain depth labels include basic layer depth, method layer depth, application layer depth, and cross-layer depth. Specifically, the core connotations of the four-level labels are defined based on the logical hierarchy of academic research. The basic layer depth focuses on fundamental domain concepts and theoretical frameworks, such as the basic principles of deep learning and the definition of Bayes' theorem; the method layer depth emphasizes specific technical methods and model improvements, such as BERT-based semantic matching methods and feature selection optimization using random forests; the application layer depth focuses on the practical application of technologies in specific scenarios, such as the application of random forests in medical image classification and the practice of semantic recommendation in academic literature retrieval; the cross-layer depth emphasizes cross-domain integration and innovation, such as optimization methods combining deep learning and quantum computing, and the cross-application of academic recommendation and natural language processing. For each label, typical keywords and key sentences at that level are collected, such as basic layer keywords: definition, principle, framework; method layer keywords: improvement, optimization, proposal, etc. These are encoded using an embedded model to form a corresponding multi-level word and phrase vector set, which serves as the benchmark vector for matching. At the same time, each level label is associated with a corresponding value orientation. The basic level is deeply associated with the value dimension of concept definition, the methodological level is deeply associated with the value dimension of innovation points and methodological models, the application level is deeply associated with the value dimension of data and metrics, and the cross-level is deeply associated with the value dimension of theoretical principles and innovation points, providing a hierarchical adaptation basis for subsequent value assessment.
[0028] Step S130: Cosine similarity is introduced to calculate the cosine similarity between the multi-level word and sentence vector set and the keyword and sentence vector set, outputting the domain depth label for each document within its respective domain. Specifically, for each document's keyword and sentence vector set, cosine similarity is calculated with the word and sentence vector sets of the basic layer, method layer, application layer, and cross layer, respectively. The cosine similarity value ranges from 0 to 1, with values closer to 1 indicating greater semantic similarity. A normalization method of (1 + cosine similarity) / 2 is used to ensure the stability and comparability of the similarity results within the [0,1] interval. A similarity threshold is set, and the label with the highest similarity exceeding the threshold is selected as the domain depth label for that document. If the similarity of all labels is below the threshold, correction is made based on expert rules for the document's domain. For example, in the computer science field, documents containing only basic concept descriptions are labeled as basic layer depth by default.
[0029] Step S140 involves hierarchically indexing and storing the documents in the database based on their domain depth tags within their respective fields. Specifically, a two-tiered domain-depth index structure is adopted. First, the primary index is divided into partitions based on the primary domain to which the document belongs, such as computer science, medicine, and physics. Then, within each primary partition, secondary indexes are divided according to the depth tags of the basic layer, methodological layer, application layer, and cross-layer. The index storage uses a distributed database, storing the basic information of the document, keyword vectors, and domain depth tags in association, and establishing a vector index to support fast similarity queries. Within the secondary indexes, further sub-indexes can be created based on semantic elements, such as research questions, theoretical principles, and methodological models. An approximate nearest neighbor (ANN) vector index (HNSW) is constructed, where each sub-index independently stores the vector data of the corresponding element, while also recording the document's storage path and index location, generating an index directory for subsequent retrieval.
[0030] Step S200: Obtain the cross-domain peer literature database from the hierarchical cross-domain literature database according to the domain depth of the current literature.
[0031] Specifically, the metadata of the current document is parsed to determine its domain and domain depth label, such as methodological depth. Then, through a hierarchical index directory, documents in all other domains that are also labeled as methodological depth are traversed, and the core information of these documents is extracted to construct a cross-domain peer-level literature database. This can be achieved by introducing multi-path recall strategies, such as recall based on research question elements, recall aligned with theoretical and methodological elements, and recall constrained by data / metric elements. From documents in all other domains that are also labeled as methodological depth, multi-path retrieval is performed through element sub-indexes. After merging the recall results, the core information of these documents is extracted to construct a cross-domain peer-level literature database. For example, using the research question element vector of the current document as the query, the top 200 documents with the highest similarity are retrieved from the research question element sub-indexes of the methodological layer in each domain; using the theoretical principles and methodological model element vectors as the query, the top 150 documents with the highest similarity are retrieved from the corresponding element sub-indexes; and using data and metric element vectors as constraints, documents with a similarity lower than 0.4 are filtered out. At the same time, the literature database is deduplicated based on DOI or title normalization, and low-quality literature, such as those published in non-core journals or those with zero citations, is filtered out.
[0032] For example, if the current literature is a methodological literature in the field of recommendation systems, then the cross-domain peer literature library contains relevant literature at all methodological depths in fields such as computer vision and natural language processing.
[0033] Step S300: Obtain the multidimensional semantic elements of the current document and the multidimensional semantic element library in the cross-domain peer-level document library through the semantic element extraction template.
[0034] Specifically, the design incorporates an extraction template with seven core elements: research questions, theoretical principles, methodological models, data, metrics, software systems, and instruments / equipment. It also includes four high-value sentence categories: innovative sentences, author viewpoint sentences, research question sentences, and concept definition sentences. A large language model, such as Qwen-14B, combined with a prompting process, is used for element extraction. An example prompt template is: "Please extract research question elements from the following literature text, requiring the output of specific sentences and confidence scores: [Literature text fragment]". For the current literature, multi-dimensional semantic elements are extracted in batches using this template. Each element includes a specific text fragment, extraction confidence score, and a sentence-level importance score for high-value sentences. The above extraction process is repeated for all literature in the cross-disciplinary, same-level literature database. All elements are integrated to form a multi-dimensional semantic element database, categorized and stored by element type, such as subsets of research question elements and methodological model elements, and a mapping between elements and literature is established. In the process of extracting from cross-domain peer literature databases, quality control steps can be performed, such as filtering element fragments with confidence scores below a preset threshold; performing multiple rounds of extraction on the same document using multiple different random seeds and taking the intersection of the results; and merging fragments with cosine similarity greater than a preset threshold among elements of the same type through vector clustering, retaining the fragment with the highest confidence score and the composite weight.
[0035] Step S400: Train the value transfer weight mapping model. The value transfer weight mapping model is used to perform value transfer weight mapping of cross-domain literature to obtain the value transfer weight in each domain.
[0036] Specifically, the model learns the correlation patterns of cross-disciplinary literature value, assigning appropriate weights to semantic elements in different fields. The model employs a deep learning architecture, such as a multilayer perceptron (MLP), taking seven semantic element vectors from cross-disciplinary literature as input and outputting the value transfer weights of these seven semantic elements for each field. The training process uses citation counts and journal impact factors as primary monitoring signals, while a four-dimensional academic value score is introduced as an auxiliary monitoring signal. For each sample document, innovative sentences, author viewpoint sentences, research question sentences, and concept definition sentences are extracted. The scores for each of these four categories are calculated and summed to obtain the total academic value score. A multi-objective loss function is constructed, such as total loss = 0.6 × citation prediction loss + 0.2 × journal impact factor loss + 0.2 × total academic value score prediction loss. By optimizing this loss function, the model learns the weights of semantic elements that contribute differently to the value of the literature in different fields. For example, in the methodological layer, methodological model elements have higher weights, while in the application layer, data and metrics elements have higher weights. After training, the model can output the value transfer weights of each semantic element within the field to which the input document belongs.
[0037] For example, in the field of computer vision, the weight of the research question element in the model output is 0.2, the weight of the method model element is 0.35, and the weight of the data element is 0.25.
[0038] In one possible implementation, a value transfer weight mapping model is trained, and step S400 further includes step S410, constructing a multi-domain comparison literature sample and performing cross-domain citation count annotation on the multi-domain comparison literature sample. Specifically, the construction of the multi-domain comparison literature sample covers a target cross-domain set, such as computer science, medicine, and biology. Literature from different research directions and publication years are selected for each domain to ensure the diversity and representativeness of the sample. The cross-domain citation count annotation obtains the number of times each sample literature is cited by literature from other domains through an academic database API. For example, if a methodological literature in the computer science field is cited 5 times by literature in the medical field and 3 times by literature in the biological field, then its cross-domain citation count annotation is 8. Meanwhile, four benchmark databases were constructed for each sample document to score its academic value. Among them, the innovation benchmark database gathers representative innovative points in the field, the viewpoint benchmark database gathers typical theoretical positions and evaluation criteria sentences, the unsolved problems benchmark database gathers key scientific questions in the field, and the concept definition benchmark database gathers consensus concept definition sentences in the field. Based on the benchmark databases, the four-dimensional academic value score of the sample document is calculated and labeled. At the same time, auxiliary information such as the field type, the journal in which it was published, and the impact factor are also labeled to provide rich features for model training.
[0039] Step S420 involves extracting multidimensional semantic elements from the multi-domain comparison document samples using a semantic element extraction template to obtain multidimensional semantic element comparison samples. Specifically, the multidimensional semantic element extraction template designed in step S300 is used to extract elements from each multi-domain comparison document sample. The extracted content includes seven types of core semantic elements and four types of high-value sentences. For each element, a text fragment is extracted and encoded into a vector using an embedding model, while retaining the extraction confidence level. The semantic element vectors of all samples are categorized and organized by domain to form a multidimensional semantic element comparison sample set. Each sample contains a triplet structure of domain label - semantic element vector - cross-domain citation count label.
[0040] For example, the semantic element comparison sample of a certain sample in the medical field is (medicine, [research question vector, method model vector, ...], 6), which represents the semantic element vector of the literature in the medical field and the number of cross-domain citations (6 times).
[0041] Step S430: Initialize value transfer weights. Calculate the loss data for the multi-dimensional semantic element comparison samples based on the initialized value transfer weights. The loss data includes cross-domain alignment loss, annotation prediction loss, and cross-domain transferability regularization loss. Specifically, the value transfer weights are initialized using the Xavier initialization method, assigning initial weights to each semantic element in each domain, with a total weight sum of 1. The cross-domain alignment loss uses cosine distance to calculate the difference between semantic element vectors of the same type across different domains. The formula is: Cross-domain alignment loss = 1 - Mean cosine similarity, ensuring semantic consistency of the same type of elements across domains. The annotation prediction loss uses mean squared error loss, simultaneously calculating the difference between the cross-domain citation count predicted by the model based on the semantic element weights and the actual annotation count, as well as the difference between the predicted total academic value score and the actual total score. The weighted sum of these two is used as the final annotation prediction loss. The cross-domain transferability regularization loss uses L2 regularization to constrain the range of value transfer weights and avoid overfitting. The formula is: Regularization loss = λ × Sum of squared weights, where λ is the regularization coefficient. The total loss data is a weighted sum of three types of losses, such as cross-domain alignment loss with a weight of 0.3, annotation prediction loss with a weight of 0.5, and regularization loss with a weight of 0.2.
[0042] For example, under a certain domain initial weight, the cross-domain alignment loss is 0.2, the annotation prediction loss is 0.3, and the regularization loss is 0.05. Then the total loss is 0.2×0.3+0.3×0.5+0.05×0.2=0.22.
[0043] Step S440: Based on the loss data, perform gradient optimization on the initial value transfer weights until the value transfer weights mapped to each domain are obtained, generating a value transfer weight mapping model. Specifically, the stochastic gradient descent (SGD) optimization algorithm is used, with a learning rate of 0.001 and a batch size of 32. In each iteration, the gradient of the total loss function with respect to each value transfer weight is calculated, and the weight values are adjusted according to the gradient direction; for example, the weights are decreased when the gradient is positive and increased when the gradient is negative. During the iteration process, the loss change of the validation set is monitored. The iteration stops when the validation set loss does not decrease for five consecutive rounds. Finally, the optimal value transfer weights of each semantic element in each domain are saved, and the weights of the four-dimensional academic value dimension corresponding to each domain are recorded to form a value transfer weight mapping model. The model is stored in dictionary form, such as {Computer Science Domain: {Research Question: 0.2, Method Model: 0.35……}, Medical Domain: {Research Question: 0.25, Data: 0.3……}}.
[0044] Step S500: Align the multidimensional semantic elements and the multidimensional semantic element library with the migration weights according to the value migration weights in each domain, and output the identified documents in each domain whose first document value is greater than the preset document value threshold. The first document value is the document value obtained based on the multidimensional semantic elements.
[0045] Specifically, the process involves obtaining the multidimensional semantic element vectors of the current document and the multidimensional semantic element vectors of a cross-domain peer-level document library. For each target domain, a value transfer weight mapping model is invoked to obtain the value transfer weight for that domain. The similarity between each semantic element vector of the current document and the corresponding element vectors of similar types in the element library is calculated. A weighted similarity sum is then calculated based on the value transfer weights. This weighted similarity sum is then fused with the four-dimensional academic value score using the following formula: First Document Value = ω × Weighted Semantic Element Similarity × 100 + (1-ω) × Total Academic Value Score, where ω is the weight and serves as the first document value. A preset document value threshold is set, and documents with a first document value greater than the threshold are selected as identifying documents. Identifying documents includes key criteria for value assessment, such as which semantic elements contribute the most.
[0046] In one possible implementation, the multidimensional semantic elements and the multidimensional semantic element library are aligned by transfer weights according to the value transfer weights under each domain. Step S500 further includes step S510, configuring a semantic vector mapping space, and inputting the multidimensional semantic elements and the multidimensional semantic element library into the semantic vector mapping space respectively to obtain multidimensional semantic element vectors and a multidimensional semantic element vector library. Specifically, the semantic vector mapping space adopts a domain-adaptive embedding space, and adversarial training is used to make the semantic element vectors from different domains comparable in the same space. The multidimensional semantic elements of the current document and text fragments from the multidimensional semantic element library are respectively input into a pre-trained embedding model, such as Qwen3-Embedding, for encoding to obtain initial vectors. The initial vectors are then input into the semantic vector mapping space, and through domain-adaptive transformation, such as minimizing domain differences through a domain discriminator, multidimensional semantic element vectors and a multidimensional semantic element vector library in a unified space are obtained, ensuring the semantic comparability of cross-domain elements.
[0047] For example, when the method model element vectors in the computer science field and the method model element vectors in the medical field are input into the mapping space, the distance between the two in the space can reflect their substantial similarity at the method level, rather than spurious differences caused by domain differences.
[0048] Step S520: Align the multidimensional semantic element vectors with the multidimensional semantic element vector library, evaluate the document value according to the value transfer weights under each domain, and output the first document value evaluation result. Specifically, the alignment of the multidimensional semantic element vectors with the element library vectors adopts a per-element type matching method, that is, the research question vector of the current document is only similar to the research question vectors of documents in the element library, avoiding cross-interference between different types of elements. The similarity calculation uses cosine similarity to obtain the similarity score for each element type. Then, according to the value transfer weights of each domain, the similarity scores of each element type are weighted and summed, and then multiplied by 100 to convert to a first document value of 0-100. The evaluation result includes the total value score of each document, the similarity score of each element type, and the weight contribution ratio.
[0049] Step S530: Utilize the first document value assessment results to filter identified documents in each field whose first document value exceeds a preset document value threshold. Specifically, the preset document value threshold is set according to field characteristics and user needs, and a dynamic threshold strategy can be adopted, such as setting a threshold for the top 30% of document value scores in each field. For cross-field peer-level document databases in each field, sort documents in descending order of first document value, and filter documents with scores greater than the corresponding threshold as identified documents. Simultaneously, record the core information of the identified documents, including title, DOI, value score, key contribution elements, etc., and generate a screening report explaining the value advantages of each identified document, such as high similarity of method / model elements and a weighted contribution ratio of 40%. Construct a structured JSON-formatted explanation chain for each identified document, including relevance evidence such as element matching pairs, original text fragments, and similarity, as well as value evidence such as high-value sentences, chapter positions, and contribution levels.
[0050] In one possible implementation, the first document value in each domain is output. Step S500 further includes step S540, which involves performing migration weight alignment on the multidimensional semantic elements and the multidimensional semantic element database, and calculating the cosine similarity of the semantic elements. Specifically, the migration weight alignment first clarifies the value priority of each semantic element, based on the weights output by the value migration weight mapping model. For example, the weight of the method model element (0.35) is higher than the weight of the instrument and equipment element (0.03), ensuring that the matching accuracy of high-weight elements is prioritized. For each target document in the cross-domain same-level document database, point-to-point matching is performed according to element type. For example, the research question element vector of the current document is only similar to the research question element vector of the target document, and the method model element vector is only similar to the method model element vector of the target document. This process is repeated for all semantic elements. The similarity calculation uses the normalized cosine similarity function, with a value range of 0-1. The closer the value is to 1, the higher the semantic fit of the element type. For each target document, after completing the matching of all elements, the weighted similarity is calculated by combining the confidence value of the elements with the similarity value multiplied by the confidence value. The associated data of current document-target document-element type-weighted similarity score is recorded to form the element similarity set of a single target document. For multiple target documents, multiple independent sets of similarity data are formed.
[0051] Step S550: Weight the cosine similarity of semantic elements according to the value transfer weight under each domain to obtain the first document value of each document under each domain. Specifically, for each target document in the cross-domain same-level document database, the value calculation process is executed. First, the value transfer weight mapping model is called to obtain the specific weights of various semantic elements under the current domain. Then, the similarity scores of various elements corresponding to the target document are extracted. For each type of element, the product of the element weighted similarity and the corresponding weight is calculated. The contribution values of all elements are added together to obtain the similarity weighted sum of the target document. The similarity weighted sum is integrated with the four-dimensional academic value total score and multiplied by 100 to convert it into a quantitative score of 0-100, which is the first document value of the target document. The above calculation process is repeated for each document in the database to obtain its own first document value.
[0052] In one possible implementation, after obtaining the multidimensional semantic elements of the current document and the multidimensional semantic element library in the cross-domain peer-level document library through a semantic element extraction template, the method further includes step S600, generating multidimensional extended sentences composed of the multidimensional semantic elements of the current document. Specifically, for each type of multidimensional semantic element in the current document, a large language model, such as Qwen-7B, is used to expand the sentences. The expansion rules include: for research question elements, supplementing the problem background and research significance; for method model elements, supplementing the method principle and implementation steps; for data elements, supplementing the data type and acquisition method, etc. During the expansion process, core information related to innovation points, author viewpoints, research questions, and concept definitions is retained first. Finally, all expanded sentences are integrated to form a set of multidimensional extended sentences, with each sentence associated with a corresponding semantic element type, confidence level, and sentence-level importance score.
[0053] Step S700: Generate a multidimensional extended statement library based on the multidimensional semantic element library. Specifically, using the same extension rules and large language model as in step S600, each element in the multidimensional semantic element library of the cross-domain peer-level document library is extended. Consistency of element types is maintained during the extension process; for example, research question elements in the element library are only extended into research question-type extended statements. The extended statements of all documents are categorized and stored by domain and element type to construct a multidimensional extended statement library. Each record in the library contains the association information of document ID, element type, extended statement, and original element fragment. Simultaneously, duplicate extended statements are deduplicated; for example, only one identical extended statement is retained, and multiple document IDs are associated to improve subsequent comparison efficiency. A quality control step is performed, conducting multiple rounds of consistency checks on the extended statements, such as generating extended statements based on two different prompt words and taking the intersection; 5% of the extended statements are selected for manual correction, and the accuracy rate is calculated. If the accuracy rate is below 80%, the extension prompt template is adjusted to ensure the accuracy and semantic integrity of the extended statements.
[0054] Step S800: Align the multidimensional extended statements and the multidimensional extended statement library with migration weights according to the value migration weights of each domain, and output the identified documents whose second document value is greater than a preset document value threshold in each domain. The second document value is the document value obtained based on the multidimensional extended statements. Specifically, the multidimensional extended statements and statements in the multidimensional extended statement library are encoded into vectors through an embedding model and input into the semantic vector mapping space configured in step S510 to obtain extended statement vectors in a unified space. According to the value migration weights of each domain, the cosine similarity of the extended statement vectors of each element type is calculated to obtain the extended statement similarity score of each element type. Using the same weighted summation method as the first document value, the weighted sum of extended statement similarity is calculated. High-value sentences are extracted based on the extended statements. Referring to the four-dimensional academic value scoring system, the total academic value score of the extended statements is calculated. Among them, the innovation point dimension score is based on the similarity between the innovation description in the extended statement and the innovation benchmark library, the viewpoint dimension score is based on the consistency and self-consistency of viewpoints, the research question dimension score is based on the necessity and clarity of the question, and the concept definition dimension score is based on consensus consistency and contribution. The second document value is obtained by fusion. A preset threshold, identical to the first document value, is set to filter out documents with a second document value greater than the threshold. The filtering results from the first and second document values are then integrated using an intersection strategy, where both value thresholds are simultaneously met to determine the final identified documents, thus improving the reliability of the filtering results. For example, a document with a first document value of 75 and a second document value of 72, both exceeding the threshold of 60, is therefore identified as a final identified document.
[0055] In one possible implementation, a cross-domain peer-level literature library is obtained from the hierarchical cross-domain literature library based on the domain depth of the current literature. The method further includes step S900: when the number of identified documents output from the peer-level cross-domain literature library is less than the number required by the user, a cross-domain neighbor-level literature library is obtained from the cross-domain literature library. Specifically, a threshold for the number of identified documents required by the user is set, such as 20 documents. The number of identified documents output from the peer-level cross-domain literature library is counted. If the number is insufficient, such as only 12 documents, a literature library at the neighbor-level domain depth is obtained. The neighbor level is defined as follows: the neighbor level of the basic layer is the method layer; the neighbor level of the method layer is the basic layer and the application layer; the neighbor level of the application layer is the method layer and the cross layer; and the neighbor level of the cross layer is the application layer. The method for obtaining the neighbor-level literature library is the same as that of the peer-level literature library. Literature across domains and corresponding neighbor-level depths needs to be filtered, and a multi-dimensional semantic element library and extended sentence library for the neighbor-level literature are constructed. Simultaneously, the depth tags of the neighbor-level literature are recorded, such as method layer - neighbor application layer, for differentiation.
[0056] Step S1000 involves using the identified documents from the cross-domain neighboring literature database as incremental output. Specifically, for the cross-domain neighboring literature database, steps S300-S800 are repeated to calculate the first and second document values of neighboring documents, and identifying documents with values greater than a preset threshold. These documents are then sorted in descending order of their document value scores, and the top N documents (N = the number required by the user - the number of identified documents at the same level) are selected as incremental documents. During incremental output, the depth level and value score of the documents are clearly marked, and semantic connections between neighboring documents and the current document are supplemented. Finally, the identified documents at the same level and the incremental identified documents are integrated to form a complete recommendation list output to the user.
[0057] For example, if a user needs 20 labeled documents, 12 documents will be output from the same level, 15 documents from the neighboring level document library will be selected to meet the criteria, and the top 8 documents will be selected as incremental outputs based on their value scores, resulting in a final output of 20 labeled documents.
[0058] This application's embodiments construct a hierarchical cross-domain literature library by performing domain-depth annotation and hierarchical indexing and storage on multi-domain literature. Based on the domain depth of the current literature, cross-domain peer-level literature is selected from the library to form a dedicated literature library. Through semantic element extraction templates, multi-dimensional semantic elements of the current literature and the multi-dimensional semantic element library of the cross-domain peer-level literature library are extracted respectively. A value transfer weight mapping model is trained to obtain the value transfer weight corresponding to each domain. The two types of semantic elements are aligned according to the transfer weight, and the identification literature with the first literature value exceeding a preset threshold is calculated and selected. This completes the evaluation and selection of cross-domain high-value literature. These technical means solve the technical problems of high sample dependence, poor cross-domain adaptability, and shallow semantic understanding in existing literature value evaluation, and achieve the technical effect of accurate selection and efficient recommendation of cross-domain high-value academic literature.
[0059] In the above text, refer to Figure 1 This paper describes in detail a cross-domain document value assessment method based on semantic element modeling according to embodiments of the present invention. Next, reference will be made to... Figure 2 This invention describes a cross-domain document value assessment system based on semantic element modeling according to an embodiment of the present invention.
[0060] The cross-domain literature value assessment system based on semantic element modeling according to embodiments of the present invention addresses the technical problems of high sample dependence, poor cross-domain adaptability, and shallow semantic understanding in existing literature value assessment methods, achieving the technical effect of accurate screening and efficient recommendation of high-value academic literature across domains. The cross-domain literature value assessment system based on semantic element modeling includes: a hierarchical cross-domain literature database construction module 10, a cross-domain peer-level literature database acquisition module 20, a multi-dimensional semantic element acquisition module 30, a value transfer weight mapping model training module 40, and an identified literature output module 50.
[0061] The hierarchical cross-domain literature database construction module 10 is used to construct a hierarchical cross-domain literature database. The hierarchical cross-domain literature database is indexed and stored hierarchically according to the domain depth of the literature in the database by performing domain depth annotation on the literature in the database. The cross-domain peer-level literature database acquisition module 20 is used to obtain the cross-domain peer-level literature database from the hierarchical cross-domain literature database according to the domain depth of the current literature. The multi-dimensional semantic element acquisition module 30 is used to obtain the multi-dimensional semantic elements of the current literature and the multi-dimensional semantic element database in the cross-domain peer-level literature database through the semantic element extraction template. The value transfer weight mapping model training module 40 is used to train the value transfer weight mapping model. The value transfer weight mapping model is used to perform value transfer weight mapping of cross-domain literature to obtain the value transfer weight under each domain. The identified literature output module 50 is used to align the multi-dimensional semantic elements and the multi-dimensional semantic element database according to the value transfer weight under each domain, and output the identified literature under each domain whose first literature value is greater than a preset literature value threshold.
[0062] The detailed configuration of the hierarchical cross-domain literature database construction module 10 is explained as follows: As mentioned above, the literature in the database is hierarchically indexed and stored according to the labeled domain depth. The hierarchical cross-domain literature database construction module 10 may further include: a vector extraction unit for extracting keyword vectors and key sentence vectors from the literature in the database to obtain a set of keyword and sentence vectors; a multi-level domain depth label definition unit for defining multi-level domain depth labels and corresponding multi-level word and sentence vector sets, wherein the multi-level domain depth labels include basic layer depth, method layer depth, application layer depth, and cross layer depth; a cosine similarity calculation unit for introducing cosine similarity to calculate the cosine similarity between the multi-level word and sentence vector set and the keyword and sentence vector set, and outputting the domain depth label of each literature in its respective domain; and a hierarchical index storage unit for hierarchically indexing and storing the literature in the database according to the domain depth label of each literature in its respective domain.
[0063] The detailed description of the specific configuration of the document identification output module 50 is explained as follows: As mentioned above, the multidimensional semantic elements and the multidimensional semantic element library are aligned according to the value transfer weight under each domain. The document identification output module 50 may further include: a semantic vector mapping unit for configuring a semantic vector mapping space, inputting the multidimensional semantic elements and the multidimensional semantic element library into the semantic vector mapping space respectively to obtain multidimensional semantic element vectors and a multidimensional semantic element vector library; a document value assessment unit for aligning the multidimensional semantic element vectors and the multidimensional semantic element vector library, performing document value assessment according to the value transfer weight under each domain, and outputting a first document value assessment result; and a filtering unit for using the first document value assessment result to filter document identification documents whose first document value under each domain is greater than a preset document value threshold.
[0064] The detailed description of the specific configuration of the value transfer weight mapping model training module 40 is explained below: As mentioned above, the value transfer weight mapping model training module 40 may further include: a cross-domain citation count annotation unit for constructing multi-domain comparison document samples and annotating the cross-domain citation counts of the multi-domain comparison document samples; a multi-dimensional semantic element extraction unit for extracting multi-dimensional semantic elements from the multi-domain comparison document samples using a semantic element extraction template to obtain multi-dimensional semantic element comparison samples; a loss data calculation unit for initializing value transfer weights and calculating the loss data of the multi-dimensional semantic element comparison samples based on the initialized value transfer weights; and a gradient optimization unit for performing gradient optimization on the initialized value transfer weights based on the loss data until the value transfer weights mapped to each domain are obtained, thereby generating a value transfer weight mapping model.
[0065] The loss data calculation unit for the multidimensional semantic element comparison samples is calculated based on the initial value migration weight. The loss data calculation unit may further include: the loss data includes cross-domain alignment loss, annotation prediction loss, and cross-domain migration regularization loss of semantic elements.
[0066] The system, after obtaining the multidimensional semantic elements of the current document and the multidimensional semantic element library in the cross-domain peer-level document library through the semantic element extraction template, may further include: a multidimensional extended statement generation module for generating multidimensional extended statements composed of the multidimensional semantic elements of the current document; a multidimensional extended statement library generation module for generating a multidimensional extended statement library based on the multidimensional semantic element library; and a migration weight alignment module for aligning the multidimensional extended statements and the multidimensional extended statement library according to the value migration weight under each domain, and outputting identified documents whose second document value is greater than a preset document value threshold under each domain.
[0067] The system may further include: the first document value is a document value obtained based on multidimensional semantic elements, and the second document value is a document value obtained based on multidimensional extended sentences.
[0068] The system may further include: a cross-domain peer literature library obtained from the hierarchical cross-domain literature library based on the current literature's domain depth; and an incremental output module used to output the identified literature from the cross-domain peer literature library as an incremental output.
[0069] Specifically, the document output module 50, which outputs the first document value in each domain, may further include: a semantic element cosine similarity calculation unit for aligning the multidimensional semantic elements and the multidimensional semantic element library by migration weights and calculating the semantic element cosine similarity; and a weighting unit for weighting the semantic element cosine similarity according to the value migration weights in each domain to obtain the first document value of each document in each domain.
[0070] The cross-domain literature value assessment system based on semantic element modeling provided in this invention can execute the cross-domain literature value assessment method based on semantic element modeling provided in any embodiment of this invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0071] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.
[0072] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A cross-domain document value assessment method based on semantic element modeling, characterized in that, The method includes: A hierarchical cross-domain literature database is constructed, wherein the literature in the database is labeled with domain depth and indexed and stored hierarchically according to the labeled domain depth. Based on the current domain depth of the literature, obtain the cross-domain peer literature library from the hierarchical cross-domain literature library; The multidimensional semantic elements of the current document are obtained through the semantic element extraction template, as well as the multidimensional semantic element library in the cross-domain peer document library. A value transfer weight mapping model is trained, which is used to perform value transfer weight mapping of cross-domain literature to obtain the value transfer weight in each domain. The multidimensional semantic elements and the multidimensional semantic element library are aligned according to the value transfer weight in each domain, and the identified documents whose first document value is greater than the preset document value threshold in each domain are output. The literature in the database is hierarchically indexed and stored according to the depth of the labeled domain. Methods include: Keyword vectors and key sentence vectors are extracted from the documents in the database to obtain a set of keyword and sentence vectors; Define multi-level domain depth labels and corresponding multi-level word and sentence vector sets, wherein the multi-level domain depth labels include base layer depth, method layer depth, application layer depth and cross layer depth; Cosine similarity is introduced to calculate the cosine similarity between the multi-level word and sentence vector set and the keyword and sentence vector set, and the domain depth label of each document in its respective field is output. The documents in the library are hierarchically indexed and stored according to the domain depth tag of each document in its respective field; Methods for training value transfer weight mapping models include: Construct a multi-domain comparison literature sample, and perform cross-domain citation count annotation on the multi-domain comparison literature sample; Multidimensional semantic elements are extracted from the multi-domain comparison document samples using a semantic element extraction template to obtain multidimensional semantic element comparison samples. Initialize value transfer weights, and calculate the loss data of the multidimensional semantic element comparison samples based on the initial value transfer weights; Based on the loss data, the initial value migration weights are gradient optimized until the value migration weights mapped to each domain are obtained, thus generating a value migration weight mapping model.
2. The cross-domain document value assessment method based on semantic element modeling as described in claim 1, characterized in that, The method further includes aligning the multidimensional semantic elements and the multidimensional semantic element library according to the value transfer weights under each domain, and performing the following: Configure a semantic vector mapping space, and input the multidimensional semantic elements and the multidimensional semantic element library into the semantic vector mapping space respectively to obtain multidimensional semantic element vectors and multidimensional semantic element vector library; Align the multidimensional semantic element vectors and the multidimensional semantic element vector library, evaluate the value of the documents according to the value transfer weights under each domain, and output the first document value evaluation result. The first document value assessment results are used to screen identified documents in each field whose first document value is greater than a preset document value threshold.
3. The cross-domain document value assessment method based on semantic element modeling as described in claim 1, characterized in that, The loss data of the multidimensional semantic element comparison samples is calculated based on the initial value transfer weights. The loss data includes cross-domain alignment loss of semantic elements, annotation prediction loss, and cross-domain mobility regularization loss.
4. The cross-domain document value assessment method based on semantic element modeling as described in claim 1, characterized in that, After obtaining the multidimensional semantic elements of the current document through the semantic element extraction template, and the multidimensional semantic element library in the cross-domain peer-level document library, the method further includes: Generate multidimensional extended sentences composed of multidimensional semantic elements of the current document; Generate a multidimensional extended statement library based on the multidimensional semantic element library; The multidimensional extended statements and the multidimensional extended statement library are aligned according to the value migration weight in each domain, and the identified documents in each domain whose second document value is greater than a preset document value threshold are output.
5. The cross-domain document value assessment method based on semantic element modeling as described in claim 4, characterized in that, The first document value is the document value obtained based on multidimensional semantic elements, and the second document value is the document value obtained based on multidimensional extended sentences.
6. The cross-domain document value assessment method based on semantic element modeling as described in claim 1, characterized in that, The method further includes obtaining a cross-domain peer-level literature database from the hierarchical cross-domain literature database based on the current literature's domain depth. When the number of identified documents output based on the cross-domain peer-level document library is less than the number required by the user, the cross-domain neighbor-level document library is obtained from the cross-domain document library; The identified documents output from the cross-domain neighboring literature database are used as incremental outputs.
7. The cross-domain document value assessment method based on semantic element modeling as described in claim 1, characterized in that, The methods for outputting the value of the first literature in each field include: The multidimensional semantic elements and the multidimensional semantic element library are aligned by transfer weights, and the cosine similarity of the semantic elements is calculated. The cosine similarity of semantic elements is weighted according to the value transfer weight in each domain to obtain the first document value of each document in each domain.
8. A cross-domain document value assessment system based on semantic element modeling, characterized in that, The system is used to implement the cross-domain document value assessment method based on semantic element modeling as described in any one of claims 1-7, and the system comprises: A hierarchical cross-domain literature database construction module is used to construct a hierarchical cross-domain literature database. The hierarchical cross-domain literature database performs domain depth annotation on the literature in the database and stores the literature in the database hierarchically according to the annotation domain depth. The cross-domain peer literature database acquisition module is used to acquire cross-domain peer literature databases from the hierarchical cross-domain literature databases based on the domain depth of the current literature. The multidimensional semantic element acquisition module is used to acquire the multidimensional semantic elements of the current document and the multidimensional semantic element library in the cross-domain peer-level document library through the semantic element extraction template. The value transfer weight mapping model training module is used to train the value transfer weight mapping model, which is used to perform value transfer weight mapping of cross-domain literature to obtain the value transfer weight in each domain. The document identification output module is used to align the multidimensional semantic elements and the multidimensional semantic element library according to the value transfer weight under each domain, and output the identified documents whose first document value under each domain is greater than a preset document value threshold.
Citation Information
Patent Citations
Literature reference value evaluation system and method based on big data
CN107273431A
Scientific paper academic value measurement method based on semantic fusion and reference behavior analysis
CN120316274A