Cross-domain knowledge migration method and system for academic literature

By acquiring and analyzing multi-field academic literature, building a cross-field knowledge transfer channel, solving the problem of cross-field knowledge transfer, realizing the automation and accurate transfer of knowledge, and improving the efficiency and innovation of academic research.

CN120338077AActive Publication Date: 2025-07-18JIEHELIX (SHANGHAI) MEDICAL TECH CO LTD

Patent Information

Application Number
CN202510813756.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

It is difficult for existing technologies to effectively realize cross-domain knowledge transfer, resulting in waste of knowledge resources and inefficient research in academic research.

Method used

By obtaining a collection of multi-field academic literature, the domain feature representation of each document unit is extracted, the cross-domain knowledge transfer channel is built, the correlation mapping relationship between the feature representations of different disciplines and fields is established, and the effectiveness of the knowledge transfer results is verified.

Benefits of technology

It has achieved automation and accurate transfer of cross-domain knowledge, and improved the efficiency and innovation of cross-integration of disciplines and academic research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338077A_ABST
    Figure CN120338077A_ABST
Patent Text Reader

Abstract

The invention provides a cross-domain knowledge migration method and system for academic literatures, and aims to solve the problem that cross-domain knowledge migration is difficult to effectively realize in the prior art. The method comprises the following steps: firstly, acquiring a multi-field academic literature set comprising a plurality of subject categories; then, field feature extraction processing is carried out on the multi-field academic literature set, and field feature representation of each literature unit is obtained; then, constructing a cross-domain knowledge migration channel based on the domain feature representations, and establishing an association mapping relationship between the domain feature representations of different subjects; utilizing the cross-domain knowledge migration channel to migrate knowledge of the source subject literature unit to the target subject literature unit, and generating a knowledge migration result; and finally, validity verification processing is carried out on the knowledge migration result, migration validity evaluation information is generated, automatic and accurate migration of cross-domain knowledge is realized, and subject cross fusion and academic research innovation are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology. Specifically, it relates to a cross-domain knowledge transfer method and system for academic literature. Background Art

[0002] In today's academic research field, with the continuous deepening of interdisciplinary integration, cross-domain knowledge transfer has become an important research direction. Although there are professional differences between different disciplinary fields, there are often many common knowledge and methods. However, most traditional academic research methods are limited to a single disciplinary field, making it difficult to effectively explore and utilize the knowledge connections between interdisciplinary fields, resulting in a waste of knowledge resources and low research efficiency.

[0003] Although natural language processing and machine learning technologies have made certain progress in the field of academic literature analysis in recent years, most existing methods focus on the superficial analysis of literature content or knowledge mining within the same field, and fail to effectively achieve cross-domain knowledge transfer. Summary of the Invention

[0004] In view of the problems mentioned above, in combination with the first aspect of the present invention, embodiments of the present invention provide a cross-domain knowledge transfer method for academic literature. The method includes: Obtain a multi-domain academic literature set containing multiple disciplinary categories. The multi-domain academic literature set is composed of literature units corresponding to different disciplines. Each literature unit contains text content and domain identification information of the discipline to which it belongs; Perform domain feature extraction processing on the multi-domain academic literature set to obtain the domain feature representation of each literature unit. The domain feature representation includes semantic features reflecting the core content of the literature unit and structural features reflecting the structural association of the literature unit; Construct a cross-domain knowledge transfer channel based on the domain feature representation. The cross-domain knowledge transfer channel is used to establish an association mapping relationship between domain feature representations of different disciplines; Use the cross-domain knowledge transfer channel to transfer the knowledge of the source discipline literature unit to the target discipline literature unit, and generate a knowledge transfer result of the target discipline literature unit; Perform validity verification processing on the knowledge transfer result to generate transfer validity evaluation information indicating the reliability degree of knowledge transfer.

[0005] In another aspect, embodiments of the present invention further provide a cross-domain knowledge transfer system for academic literature, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0006] Based on the above aspects, through the comprehensive application of technical means such as domain feature extraction, cross-domain knowledge transfer channel construction, and knowledge transfer effectiveness verification, the automated and cross-domain transfer of knowledge in multi-domain academic literature is realized. It can not only accurately extract the domain feature representations of each literature unit, covering semantic features and structural features, but also construct cross-domain knowledge transfer channels, effectively establish the correlation mapping relationship between the domain feature representations of different disciplinary fields, and achieve cross-domain transfer of knowledge; finally, through the effectiveness verification process of the knowledge transfer results, evaluation information indicating the reliability of knowledge transfer is generated to ensure the accuracy of the transferred knowledge, significantly improving the automation level and accuracy of cross-domain knowledge transfer, and contributing to promoting interdisciplinary integration, enhancing the efficiency and innovation of academic research. Brief Description of the Drawings

[0007] Figure 1 It is a schematic flowchart of the execution process of the cross-domain knowledge transfer method for academic literature provided by an embodiment of the present invention.

[0008] Figure 2 It is a schematic diagram of the exemplary hardware and software components of the cross-domain knowledge transfer system for academic literature provided by an embodiment of the present invention. Detailed Embodiments

[0009] The present invention will be specifically described below in conjunction with the accompanying drawings of the specification. Figure 1 It is a schematic flowchart of the cross-domain knowledge transfer method for academic literature provided by an embodiment of the present invention. The cross-domain knowledge transfer method for academic literature will be introduced in detail below.

[0010] Step S110: Obtain a multi-domain academic literature set containing multiple disciplinary categories. The multi-domain academic literature set is composed of literature units corresponding to different disciplines, and each literature unit contains text content and domain identification information of the discipline to which it belongs.

[0011] In this embodiment, in order to carry out cross - domain knowledge transfer work for academic literature, it is first necessary to construct a multi - domain academic literature collection. This multi - domain academic literature collection should cover multiple disciplinary categories, and each discipline corresponds to a number of literature units. For example, in a multi - domain academic literature collection involving medicine, computer science, and sociology, the literature units under the medical discipline may be academic papers on disease treatment, drug research and development, etc. The text content includes specific case analyses, experimental data, etc., and carries the domain identification information of the medical discipline; the literature units under the computer science discipline may be research reports on artificial intelligence algorithms, software development, etc. The text content includes algorithm design, code implementation, etc., and also carries the domain identification information of the computer science discipline; the literature units under the sociology discipline may be academic works on social phenomenon analysis, group behavior research, etc. The text content includes social survey data, case analyses, etc., and also carries the domain identification information of the sociology discipline.

[0012] These literature units can be obtained through various channels, such as searching academic databases. Professional academic databases, such as CNKI, Wanfang, Web of Science, etc., can be used to search according to keywords in different disciplines. Taking the medical discipline as an example, keywords such as "cancer treatment", "cardiovascular disease research" are used for searching, and the relevant literature retrieved is screened and incorporated into the multi - domain academic literature collection. Literature units can also be obtained through channels such as the websites of academic institutions and academic conference proceedings. For the computer science discipline, obtain cutting - edge research papers from the websites of well - known academic institutions such as the Department of Computer Science at Stanford University, or collect relevant literature from the proceedings of important academic conferences such as ACM SIGKDD.

[0013] During the process of collecting literature units, it is necessary to ensure that each literature unit has accurate domain identification information of its affiliated discipline. The affiliated discipline can be determined through information such as the keywords, abstracts, and classification numbers of the literature itself. For example, if the keywords of a piece of literature include "neural network", "deep learning", etc., and the abstract mainly discusses the algorithm applications in the field of computer science, then it can be identified as a literature unit of the computer science discipline.

[0014] In the above process, to ensure the legality and compliance of data, it is necessary to strictly follow relevant laws and regulations to obtain data authorization. For the case of obtaining literature units from public academic databases, it is necessary to carefully study the usage terms and authorization agreements of the databases. Generally speaking, most public academic databases will clearly stipulate the scope, purpose, and limitations of data use. Before obtaining data, it is necessary to confirm that the cross-domain knowledge transfer research being carried out meets the usage requirements of the database. For example, some databases may allow non-commercial research use of data, but prohibit using the data for commercial profit or other unauthorized purposes. When obtaining data, it is necessary to operate in accordance with the procedures stipulated by the database, such as registering an account, filling in a description of the usage purpose, etc., to obtain legal data access rights.

[0015] For the case of obtaining literature units from channels such as academic institution websites and academic conference proceedings, it is necessary to communicate directly with the relevant institutions or copyright owners to obtain authorization for data collection. A formal request email can be sent, detailing information such as the purpose, use, and research methods of data collection. The email should clearly state that relevant laws, regulations, and ethical guidelines will be adhered to, and the privacy and intellectual property rights of the data will be protected. For example, when requesting research papers on a university's computer science department website, it is stated in the email that these papers will be used for cross-domain knowledge transfer research aimed at promoting knowledge exchange and sharing between different disciplines, and will not be used for commercial purposes or disclosed to third parties. Wait for a reply from the relevant institution or copyright owner, and after obtaining clear authorization, proceed with data collection operations.

[0016] During the data collection process, for literature units that may involve privacy-sensitive data, strict privacy protection and anti-disclosure technical measures should be taken. For medical literature units containing patient personal information, special processing should be carried out on this privacy-sensitive information during text preprocessing. Data desensitization technology can be used to replace sensitive information such as patient names, ID numbers, and addresses with anonymous identifiers. For example, replace patient names with "Patient A", "Patient B", etc., and replace ID numbers with a string of randomly generated characters. At the same time, the processed data is encrypted for storage, using advanced encryption algorithms such as the AES encryption algorithm to ensure the security of the data during storage and transmission. When accessing and using this processed data, a strict permission management mechanism is set up, and only authorized personnel can access the relevant data.

[0017] Step S120: Perform domain feature extraction processing on the multi-domain academic literature collection to obtain the domain feature representation of each literature unit, where the domain feature representation includes semantic features reflecting the core content of the literature unit and structural features reflecting the structural association of the literature unit.

[0018] After obtaining the multi-domain academic document collection, the next step is to extract domain features from each document unit in the multi-domain academic document collection to obtain a domain feature representation that can represent the document unit. The domain feature representation consists of two parts: one is the semantic feature that reflects the core content of the document unit, and the other is the structural feature that reflects the structural association of the document unit.

[0019] Semantic features can reflect the meaning of key concepts and the expression of core arguments in a document unit. For example, in a medical document, semantic features can reflect the core content such as the pathogenesis and treatment methods of a disease; in computer science documents, semantic features can reflect the principles and application scenarios of algorithms. Structural features focus on the organizational structure of a document unit and the relationship between its parts. For example, the hierarchical relationship between different chapters in medical documents, the citation relationship between diagrams and the text, etc.; the code structure of different modules in computer science documents, the relationship between references and the text, etc.

[0020] Step S121: performing text preprocessing operations on each document unit in the multi-field academic document collection, wherein the text preprocessing operations include removing non-content symbols, filtering common stop words, and splitting continuous text paragraphs into processable text segment units.

[0021] In order to more accurately extract the domain features of document units, we first need to perform text preprocessing on each document unit in the multi-domain academic document collection.

[0022] Removing non-content symbols is the first step in preprocessing. There are many non-content symbols in the text content of the document, such as punctuation marks (commas, periods, exclamation marks, etc.) and special characters (@, #, $, etc.). Although these symbols play a certain role in text expression, they do not provide key information for semantic analysis and feature extraction, but may interfere with subsequent processing. For example, in a medical document, a large number of punctuation marks will make the semantic structure of the text complicated, which is not conducive to the identification of key concepts. Therefore, these non-content symbols need to be removed to make the text content more concise. The removal of non-content symbols can be achieved through regular expressions, such as defining a regular expression pattern to match all punctuation marks and special characters, and then replacing them with an empty string.

[0023] Filtering common stop words is the second step of preprocessing. Common stop words refer to words that frequently appear in natural language but contribute little to the core semantic expression of the text, such as "de", "shi", "zai", "he", etc. In the literature, the large presence of these words will increase the complexity of text processing and may obscure key semantic information. For example, in computer science literature, the word "de" may appear in many places, but it does not help to understand the core principles of algorithms. Therefore, it is necessary to filter out these common stop words. A stop word list can be established, including common stop words, and then during the text processing, the words in the text are compared with the stop word list, and the matching stop words are removed from the text.

[0024] Splitting continuous text paragraphs into processable text fragment units is the third step of preprocessing. A literature usually consists of multiple continuous text paragraphs, which may contain different topics and information. To facilitate subsequent semantic analysis and feature extraction, it is necessary to split continuous text paragraphs into smaller, processable text fragment units. The splitting can be done according to the topic of the paragraph, the logical relationship of the sentences, etc. For example, in a medical literature, the paragraphs describing disease symptoms and the paragraphs introducing treatment methods can be split into independent text fragment units respectively. In this way, each text fragment unit can be used as a relatively independent processing object, which is more convenient for subsequent semantic analysis.

[0025] Step S122: Call the pre-trained semantic encoding model to perform context semantic analysis on the text fragment unit, and generate semantic features reflecting the core content of the literature unit. The semantic features include the context association information of key concepts in the text fragment unit and the semantic expression vector of the core argument.

[0026] After text preprocessing, processable text fragment units are obtained. Next, call the pre-trained semantic encoding model to perform context semantic analysis on these text fragment units to generate semantic features reflecting the core content of the literature unit.

[0027] The pre-trained semantic encoding model is trained on a large-scale text data, and it can learn the semantic information and context relationship of natural language. For example, in the medical field, the pre-trained semantic encoding model can learn the semantic associations between disease names, symptoms, treatment methods, etc.; in the field of computer science, it can learn the semantic relationships between algorithm names, data structures, programming languages, etc.

[0028] Step S1221: Input the text fragment unit into the word embedding layer of the semantic encoding model to generate the initial word vector representation of the words in each text fragment unit.

[0029] First, input the text segment unit into the word embedding layer of the semantic encoding model. The role of the word embedding layer is to convert each vocabulary in the text into a corresponding vector representation, that is, the initial word vector representation. In this process, the word embedding layer will map the vocabulary into a high-dimensional vector space according to the pre-trained word vector model. For example, in the text segment unit of medical literature, the vocabulary "cancer" will be mapped to a specific vector, which contains the information of the vocabulary "cancer" in the semantic space. The positions and directions of different vocabularies in the vector space reflect their semantic relationships. For example, the positions of the two vocabularies "cancer" and "tumor" in the vector space may be relatively close because they are semantically related.

[0030] Step S1222: Perform temporal context modeling processing on the initial word vector representation through the bidirectional long short-term memory network layer of the semantic encoding model to capture the semantic dependency relationships of the vocabularies in the text segment unit in the context before and after, and generate an intermediate semantic vector containing context information.

[0031] Next, input the generated initial word vector representation into the bidirectional long short-term memory (Bi-LSTM) layer of the semantic encoding model. The Bi-LSTM layer is a type of recurrent neural network that can process sequential data and capture the temporal dependency relationships between elements in the sequence. In text processing, the Bi-LSTM layer can capture the semantic dependency relationships of the vocabularies in the text segment unit in the context before and after. For example, in a sentence "This drug has a significant effect on cancer treatment", there is a semantic dependency relationship between "drug" and "cancer treatment", and this relationship can be learned through the Bi-LSTM layer. The Bi-LSTM layer will process the text sequence in two directions (forward and backward) to capture the forward and backward context information of the vocabulary respectively. Then, the information in these two directions is merged to generate an intermediate semantic vector containing context information. This intermediate semantic vector not only contains the semantic information of the vocabulary itself but also the semantic information in its context before and after.

[0032] Step S1223: Use the attention mechanism layer of the semantic encoding model to perform key information focusing processing on the intermediate semantic vector, identify the key concept vocabularies strongly related to the core content of the literature unit in the text segment unit, and generate the attention weight distribution of the key concept vocabularies.

[0033] After obtaining the intermediate semantic vector, the attention mechanism layer of the semantic encoding model is used to focus on key information. The role of the attention mechanism layer is to identify the key concept words that are strongly related to the core content of the literature unit in the text segment unit, and assign attention weights to these key concept words. For example, in a text segment unit of a medical literature, "cancer treatment", "drug efficacy", etc. may be key concept words. The attention mechanism layer will calculate the relevance of each word to the core content based on the information of the intermediate semantic vector, and then assign an attention weight to each word. The higher the relevance of a word, the greater its attention weight. In the above way, the importance of key concept words can be highlighted and the interference of irrelevant information can be reduced. Finally, an attention weight distribution of key concept words is generated, which reflects the importance of each key concept word in the text segment unit.

[0034] Step S1224: Perform weighted aggregation processing on the intermediate semantic vector according to the attention weight distribution to generate a local semantic feature reflecting the context association information of key concept words.

[0035] According to the generated attention weight distribution, perform weighted aggregation processing on the intermediate semantic vector. Specifically, multiply each intermediate semantic vector by the corresponding attention weight, and then aggregate these weighted vectors. For example, for a text segment unit containing multiple words, each word has a corresponding intermediate semantic vector and attention weight. Multiply each intermediate semantic vector by its attention weight and then sum them up to obtain a local semantic feature reflecting the context association information of key concept words. The local semantic feature not only contains the semantic information of key concept words but also contains their association information in the context. Through the above weighted aggregation processing, the key semantic information in the text segment unit can be extracted more accurately.

[0036] Step S1225: Input the local semantic feature into the fully connected layer of the semantic encoding model for dimensionality compression processing to generate a semantic expression vector with a fixed-dimensional representation, and the semantic expression vector is consistent with the local semantic feature in terms of semantic information.

[0037] Finally, the local semantic features are input into the fully connected layer of the semantic encoding model for dimensionality compression. The fully connected layer is a neural network layer that can map the input vector into a vector space with a fixed dimension. In this process, the fully connected layer performs linear transformation and non-linear activation on the local semantic features, converting them into semantic expression vectors with a fixed-dimension representation. For example, the dimension of the local semantic features may be relatively high, which is not conducive to subsequent processing and analysis. Through the fully connected layer, it is compressed into a lower fixed dimension. At the same time, the fully connected layer ensures that the semantic expression vector maintains semantic information consistency with the local semantic features, that is, the semantic expression vector can still accurately reflect the core semantic information of the text fragment unit.

[0038] Step S123: Parse the text structure information of the literature unit, where the text structure information includes the hierarchical relationship of chapter titles, the citation association relationship between figures / tables and the main text, and the disciplinary distribution relationship of reference documents.

[0039] While extracting semantic features, it is also necessary to parse the text structure information of the literature unit. The text structure information includes multiple aspects. For cross-domain knowledge transfer, the hierarchical relationship of chapter titles, the citation association relationship between figures / tables and the main text, and the disciplinary distribution relationship of reference documents are relatively important information.

[0040] The hierarchical relationship of chapter titles reflects the organizational structure of the literature and the logical level of the content. For example, in a medical literature, there may be chapter titles such as "Introduction", "Disease Overview", "Treatment Methods", "Experimental Results", "Conclusion", etc. There is a certain hierarchical relationship between these chapter titles. For example, "Treatment Methods" may be a sub-chapter under "Disease Overview". By parsing the hierarchical relationship of chapter titles, the overall framework and logical order of the literature content can be understood. The format of chapter titles (such as font size, indentation, etc.) can be identified through text parsing techniques, and the hierarchical relationship between chapter titles can be determined based on this format information.

[0041] The citation association relationship between figures / tables and the main text reflects the role of figures / tables in the literature and their association with the main text content. In medical literature, there may be some figures / tables showing data such as disease incidence and treatment effects. These figures / tables are usually cited in the main text to support the discussion in the main text. By parsing the citation association relationship between figures / tables and the main text, the information flow and mutual support relationship between figures / tables and the main text can be understood. Through text matching techniques, sentences citing figures / tables can be found in the main text, and then the figures / tables can be associated with these citation sentences.

[0042] The disciplinary distribution relationship of references reflects the research background and knowledge sources of the literature. In a computer science literature, the references may involve different sub-fields of computer science or other related disciplines (such as mathematics, physics, etc.). By analyzing the disciplinary distribution relationship of references, the knowledge citation situation of the literature in the interdisciplinary field can be understood. The references can be classified and labeled to determine the discipline to which each reference belongs, and then the number of references in different disciplines can be counted to analyze their disciplinary distribution.

[0043] Step S124: Extract structure features reflecting the structural association of the literature unit based on the text structure information, where the structure features include the hierarchical depth parameter of the chapter title, the citation frequency parameter of the figure and text, and the disciplinary concentration parameter of the references.

[0044] Extract structure features reflecting the structural association of the literature unit according to the parsed text structure information. These structure features include the hierarchical depth parameter of the chapter title, the citation frequency parameter of the figure and text, and the disciplinary concentration parameter of the references.

[0045] Step S1241: Perform hierarchical depth calculation processing on the hierarchical relationship of the chapter titles in the text structure information, count the hierarchical distance of each chapter title relative to the root directory, and generate the hierarchical depth parameter of the chapter title. The hierarchical depth parameter is used to represent the core degree of the chapter content in the literature unit.

[0046] Perform hierarchical depth calculation processing on the hierarchical relationship of the chapter titles. First, determine the root directory of the literature. Usually, the root directory can be the main title of the literature. Then, count the hierarchical distance of each chapter title relative to the root directory. For example, in a medical literature, the main title is the root directory, "Disease Overview" is a first-level chapter title with a hierarchical depth of 1; "Disease Symptoms" is a second-level chapter title under "Disease Overview" with a hierarchical depth of 2. The hierarchical depth parameter can reflect the core degree of the chapter content in the literature unit. Generally speaking, the content corresponding to the chapter title with a shallower hierarchical depth may be more core. Through the hierarchical depth calculation processing, a hierarchical depth parameter can be generated for each chapter title, and these parameters constitute the features reflecting the hierarchical structure of the chapter title.

[0047] Step S1242: Perform citation frequency statistics processing on the citation association relationship between the figures and text in the text structure information, count the number of times each figure is cited by the text paragraphs, and generate the citation frequency parameter of the figure and text. The citation frequency parameter is used to represent the closeness of the association between the figure content and the text content.

[0048] Perform citation frequency statistical processing on the citation association relationship between figures / tables and the main text. By searching for sentences citing figures / tables in the main text, count the number of times each figure / table is cited. For example, in a computer science literature, there is a figure showing the performance of an algorithm, and it is cited 5 times in the main text, then the citation frequency parameter of this figure is 5. The citation frequency parameter can represent the degree of closeness of the relationship between the content of the figure / table and the main text. The higher the citation frequency, the closer the relationship between the figure / table and the main text. By counting the citation frequency of each figure / table, generate the citation frequency parameters of the figures / tables and the main text, and these parameters constitute the characteristics reflecting the citation structure between the figures / tables and the main text.

[0049] Step S1243: Perform subject concentration calculation processing on the subject distribution relationship of the references in the text structure information, count the proportion of the number of documents belonging to the same subject in the references, and generate the subject concentration parameter of the references. The subject concentration parameter is used to represent the subject focus degree of the research content of the document unit.

[0050] Perform subject concentration calculation processing on the subject distribution relationship of the references. First, classify the references to determine the subject to which each reference belongs. Then, count the proportion of the number of references belonging to the same subject in the total number of references. For example, in a medical literature, there are a total of 20 references, and 15 of them belong to the medical subject, then the subject concentration parameter of the references of this literature is 15 / 20. The subject concentration parameter can represent the subject focus degree of the research content of the document unit. The higher the subject concentration, the more focused the research content of the document is on a certain subject. By calculating the subject concentration parameter, generate the characteristics reflecting the subject distribution structure of the references.

[0051] Step S1244: Standardize the hierarchical depth parameter, the citation frequency parameter, and the subject concentration parameter to eliminate the dimensional differences between different parameters and generate a set of structural characteristics with comparable metrics.

[0052] Since the dimensionalities of the hierarchical depth parameter, the citation frequency parameter, and the subject concentration parameter are different, in order to effectively compare and analyze these parameters, they need to be standardized. The purpose of standardization is to eliminate the dimensional differences between different parameters and make them have comparable metrics. Common standardization methods can be used, such as z-score standardization. For the hierarchical depth parameter, the citation frequency parameter, and the subject concentration parameter, calculate their mean and standard deviation respectively, then subtract the mean from each parameter value and divide by the standard deviation to obtain the standardized parameter value. In the above way, convert the hierarchical depth parameter, the citation frequency parameter, and the subject concentration parameter into numerical values with the same dimension, and generate a set of structural characteristics with comparable metrics.

[0053] Step S1245: Perform feature fusion processing on the structural feature set, and integrate the hierarchical depth parameter, the citation frequency parameter, and the subject concentration parameter into a single-dimensional structural feature by weighted summation. The structural feature retains the original semantic information of each parameter.

[0054] Perform feature fusion processing on the standardized structural feature set. To obtain a feature that can comprehensively reflect the structural association of the literature unit, it is necessary to integrate the hierarchical depth parameter, the citation frequency parameter, and the subject concentration parameter. Feature fusion can be performed by weighted summation. Assign a weight to each parameter, then multiply each parameter by its weight and add them together to obtain a single-dimensional structural feature. For example, assign weight w1 to the hierarchical depth parameter, weight w2 to the citation frequency parameter, and weight w3 to the subject concentration parameter. Multiply the hierarchical depth parameter by w1, the citation frequency parameter by w2, and the subject concentration parameter by w3, and then add them together to obtain the structural feature. When assigning weights, the importance of each parameter needs to be considered to ensure that the structural feature can retain the original semantic information of each parameter.

[0055] Step S125: Standardize the semantic feature and the structural feature to generate a domain feature representation with a unified representation form, which contains both semantic association information and structural association information.

[0056] After obtaining the semantic feature and the structural feature, in order to process and compare them in the same framework, it is necessary to standardize them. The dimensions and value ranges of the semantic feature and the structural feature may be different. Through standardization, these differences can be eliminated to generate a domain feature representation with a unified representation form. A standardization method similar to step S1244 can be used, such as z-score standardization. Calculate the mean and standard deviation of the semantic feature and the structural feature respectively, then subtract the mean of each element of the semantic feature and the structural feature, and divide by their standard deviation to obtain the standardized semantic feature and structural feature.

[0057] Next, perform a concatenation operation on the standardized semantic feature and the structural feature. Since the semantic feature and the structural feature represent different aspects of information of the literature unit respectively, concatenation can integrate these two types of information together to form a unified domain feature representation. For example, assume that the semantic feature is an m-dimensional vector and the structural feature is an n-dimensional vector. Then the concatenated domain feature representation is an m + n-dimensional vector. This domain feature representation contains both semantic association information and structural association information and can represent the features of the literature unit more comprehensively.

[0058] Step S130: Construct a cross - domain knowledge transfer channel based on the domain feature representation, where the cross - domain knowledge transfer channel is used to establish an association mapping relationship between domain feature representations of different disciplinary fields.

[0059] After obtaining the domain feature representation of each literature unit, the next step is to construct a cross - domain knowledge transfer channel based on these domain feature representations. The role of this cross - domain knowledge transfer channel is to establish an association mapping relationship between domain feature representations of different disciplinary fields, enabling knowledge to be transferred between different disciplines.

[0060] Step S131: Select cross - disciplinary literature pairs containing a common research topic from the multi - domain academic literature collection, where the cross - disciplinary literature pairs are composed of literature units belonging to different disciplines but with overlapping research topics.

[0061] First, select cross - disciplinary literature pairs from the multi - domain academic literature collection. These cross - disciplinary literature pairs need to contain a common research topic, that is, they are composed of literature units belonging to different disciplines but with overlapping research topics. For example, in the fields of medicine and biology, there may be research on gene therapy for diseases. Then, literature pairs related to gene therapy can be selected from medical literature units and biological literature units. By analyzing information such as the titles, abstracts, and keywords of literature units, literature unit pairs with overlapping research topics can be found. For example, using a text matching algorithm, compare the keywords of different literature units. If a certain number of keywords are the same or similar, it is considered that these two literature units may have a common research topic.

[0062] Step S132: Extract the domain feature representation of the source - discipline literature unit and the domain feature representation of the target - discipline literature unit in the cross - disciplinary literature pair.

[0063] After selecting the cross - disciplinary literature pair, extract the domain feature representation of the source - discipline literature unit and the domain feature representation of the target - discipline literature unit in it. The source - discipline literature unit is the source of knowledge, and the target - discipline literature unit is the target of knowledge transfer. For example, in a cross - disciplinary literature pair between medicine and computer science, if disease diagnosis knowledge in the medical field is to be transferred to an intelligent diagnosis system in the computer science field, then the medical literature unit is the source - discipline literature unit, and the computer science literature unit is the target - discipline literature unit. According to the domain feature representation of each literature unit obtained in the previous steps, directly extract the domain features of the source discipline and the target discipline from the cross - disciplinary literature pair.

[0064] Step S133: Calculate the feature similarity value between the domain feature representation of the source - discipline literature unit and the domain feature representation of the target - discipline literature unit, where the feature similarity value is calculated by the cosine similarity algorithm.

[0065] After obtaining the domain feature representations of the source discipline and the target discipline, calculate the feature similarity value between them. The cosine similarity algorithm can be used to calculate the feature similarity value. Cosine similarity is a commonly used method for calculating vector similarity, which measures the similarity between two vectors by calculating the cosine value of the angle between them. For the domain feature representation vector A of the source discipline literature unit and the domain feature representation vector B of the target discipline literature unit, first calculate the dot product of vector A and vector B, then calculate the magnitudes of vector A and vector B respectively, and finally divide the dot product by the product of the magnitudes of the two vectors to obtain the cosine similarity value. The closer the cosine similarity value is to 1, the more similar the two vectors are, that is, the more similar the domain feature representations of the source discipline and the target discipline; the closer the value is to 0, the less similar the two vectors are.

[0066] Step S134: Screen the interdisciplinary literature pairs with similarity values higher than the preset threshold according to the feature similarity value as the key mapping samples.

[0067] According to the calculated feature similarity value, screen out the interdisciplinary literature pairs with similarity values higher than the preset threshold as the key mapping samples. The preset threshold is a pre-set standard used to judge whether the similarity of the interdisciplinary literature pairs is high enough for subsequent association rule mining. For example, if the preset threshold is a set similarity value, when the feature similarity value of a certain interdisciplinary literature pair is higher than this threshold, it is selected as the key mapping sample. By screening the key mapping samples, unnecessary calculations and analyses can be reduced, and efforts can be concentrated on dealing with those interdisciplinary literature pairs with high similarity.

[0068] Step S135: Perform association rule mining on the source discipline domain feature representation and the target discipline domain feature representation of the key mapping samples to generate a set of migration rules reflecting the corresponding relationships of interdisciplinary features.

[0069] Perform association rule mining on the source discipline domain feature representation and the target discipline domain feature representation of the screened key mapping samples. The purpose of association rule mining is to find the corresponding relationships between the source discipline and the target discipline domain feature representations and generate a set of migration rules.

[0070] Step S1351: Perform feature decomposition on the source discipline domain feature representation of the key mapping samples to separate the semantic feature part and the structural feature part.

[0071] First, perform feature decomposition on the source subject field feature representation of the key mapping samples. Since the field feature representation is composed of semantic features and structural features spliced together, the semantic feature part and the structural feature part can be separated from it. According to the dimension information during the previous splicing, the field feature representation vector can be divided by dimension to obtain a semantic feature vector and a structural feature vector. For example, if the field feature representation is a vector of m + n dimensions, where the first m dimensions are semantic features and the last n dimensions are structural features, then the first m dimensions can be extracted as the semantic feature part, and the last n dimensions can be extracted as the structural feature part.

[0072] Step S1352: Perform synchronous feature decomposition on the target subject field feature representation of the key mapping samples to separate the corresponding semantic feature part and structural feature part.

[0073] Similarly, perform synchronous feature decomposition on the target subject field feature representation of the key mapping samples. According to the same method as in step S1351, the target subject field feature representation vector is separated into a semantic feature part and a structural feature part. In this way, the field feature representations of both the source subject and the target subject are decomposed into two parts: semantic features and structural features, which is convenient for subsequent correspondence analysis.

[0074] Step S1353: Establish a semantic correspondence between the semantic feature part of the source subject and the semantic feature part of the target subject. The semantic correspondence is generated by matching the feature sub-vectors representing the same research concept in the two parts.

[0075] Establish a semantic correspondence between the semantic feature part of the source subject and the semantic feature part of the target subject. The semantic correspondence can be generated by matching the feature sub-vectors representing the same research concept in the two parts. For example, in an interdisciplinary literature pair of medicine and computer science, the semantic feature part of the source subject (medicine) may have a feature sub-vector representing "disease diagnosis", and the semantic feature part of the target subject (computer science) may have a feature sub-vector representing "intelligent diagnosis algorithm". Although these two sub-vectors come from different subjects, they are both related to diagnosis, so a semantic correspondence can be established between them. A vector matching algorithm, such as the nearest neighbor algorithm, can be used to find similar feature sub-vectors in the semantic feature part of the source subject and the semantic feature part of the target subject and match them up.

[0076] Step S1354: Establish a structural correspondence between the structural feature part of the source subject and the structural feature part of the target subject. The structural correspondence is generated by matching the feature sub-vectors representing the same logical organization method in the two parts.

[0077] Establish a structural correspondence between the source subject matter structure feature part and the target subject matter structure feature part. Generate the structural correspondence by matching the feature sub-vectors representing the same logical organization method in the two parts. For example, the source subject matter (medicine) structure feature part may have a feature sub-vector representing the chapter hierarchical relationship, and the target subject matter (computer science) structure feature part may have a feature sub-vector representing the code module hierarchical structure. If the logical organization methods represented by these two sub-vectors are similar (such as both being hierarchical structures), then a structural correspondence can be established between them. Similarly, a vector matching algorithm can be used to find similar feature sub-vectors and establish the correspondence.

[0078] Step S1355: Perform rule induction processing on the semantic correspondence and the structural correspondence to generate a set of migration rules including semantic migration rules and structural migration rules. The semantic migration rules are used to guide the conversion of source subject matter semantic features to target subject matter semantic features, and the structural migration rules are used to guide the conversion of source subject matter structure features to target subject matter structure features.

[0079] Perform rule induction processing on the established semantic correspondence and structural correspondence. Rule induction is to summarize these correspondences into general rules. For the semantic correspondence, semantic migration rules are induced, and these rules are used to guide the conversion of source subject matter semantic features to target subject matter semantic features. For example, if a certain concept in the source subject matter semantic features has a correspondence with a certain concept in the target subject matter semantic features, then a rule can be induced to illustrate how to convert this concept of the source subject matter into the corresponding concept of the target subject matter during knowledge migration. For the structural correspondence, structural migration rules are induced, and these rules are used to guide the conversion of source subject matter structure features to target subject matter structure features. For example, if there is a correspondence between the chapter hierarchical structure of the source subject matter and the code module hierarchical structure of the target subject matter, then a rule can be induced to illustrate how to convert the chapter hierarchical structure of the source subject matter into the code module hierarchical structure of the target subject matter. Integrate the semantic migration rules and the structural migration rules together to form a set of migration rules.

[0080] Step S1356: Perform conflict detection processing on the set of migration rules to eliminate the contradictory descriptions between different rules and generate the final set of migration rules.

[0081] Perform conflict detection and handling on the generated set of migration rules. During the rule induction process, there may be contradictory descriptions between different rules. For example, a semantic migration rule may stipulate that a certain concept in the source discipline is converted into concept A in the target discipline, while another rule may stipulate conversion into concept B in the target discipline. Through conflict detection and handling, identify these contradictory rules and make adjustments or eliminations. The conflict points between the rules can be found by pairwise comparing the rules in the set of migration rules through a rule comparison algorithm, and then processed according to the set strategy (such as preferentially selecting more specific rules), and finally the contradictory descriptions are eliminated to generate the final set of migration rules.

[0082] Step S136: Structurally store the set of migration rules to generate a cross-domain knowledge migration channel for establishing an association mapping relationship between feature representations in different disciplinary fields.

[0083] Structurally store the final set of migration rules. Structural storage can store the set of migration rules in an orderly and queryable manner, facilitating subsequent use and management. A database can be used to store the set of migration rules, with each rule being a record in the database, recording the type of the rule (semantic migration rule or structural migration rule), source discipline feature information, target discipline feature information, etc. By structurally storing the set of migration rules, a cross-domain knowledge migration channel for establishing an association mapping relationship between feature representations in different disciplinary fields is generated. This cross-domain knowledge migration channel can serve as a bridge for knowledge migration, enabling the knowledge of the source discipline to be mapped to the target discipline through these rules.

[0084] Step S140: Use the cross-domain knowledge migration channel to transfer the knowledge of the source discipline literature unit to the target discipline literature unit to generate a knowledge migration result of the target discipline literature unit.

[0085] After constructing the cross-domain knowledge migration channel, the knowledge of the source discipline literature unit can be transferred to the target discipline literature unit using this cross-domain knowledge migration channel to generate a knowledge migration result of the target discipline literature unit.

[0086] Step S141: Select a source discipline literature unit and a target discipline literature unit that need to perform knowledge migration from the multi-domain academic literature set, and the source discipline literature unit contains the target knowledge content to be migrated.

[0087] First, select the source discipline literature unit and the target discipline literature unit that need to carry out knowledge transfer from the multi-disciplinary academic literature collection. The source discipline literature unit contains the target knowledge content to be transferred. For example, in the scenario of interdisciplinary knowledge transfer between medicine and computer science, if the knowledge about disease diagnosis in the medical field is to be transferred to the intelligent diagnosis system in the computer science field, then select the literature unit containing the knowledge related to disease diagnosis from the medical literature unit as the source discipline literature unit, and select the literature unit that needs to supplement knowledge from the computer science literature unit as the target discipline literature unit. Appropriate source discipline literature units and target discipline literature units can be screened according to specific knowledge transfer requirements through information such as the title, abstract, and keywords of the literature.

[0088] Step S142: Extract the domain feature representation of the source discipline literature unit, where the domain feature representation includes the source discipline semantic feature and the source discipline structural feature.

[0089] After selecting the source discipline literature unit, extract its domain feature representation. This domain feature representation includes the source discipline semantic feature and the source discipline structural feature. In the previous steps, domain feature extraction processing has been performed on each literature unit in the multi-disciplinary academic literature collection, so the domain feature representation of the source discipline literature unit can be directly extracted from the stored domain feature representation. The source discipline semantic feature reflects the core content semantic information of the source discipline literature unit, and the source discipline structural feature reflects the structural association information of the source discipline literature unit.

[0090] Step S143: Perform transformation processing on the source discipline semantic feature through the semantic transfer rule in the cross-domain knowledge transfer channel to generate a target semantic feature aligned with the dimension of the target discipline semantic feature.

[0091] Perform transformation processing on the source discipline semantic feature using the semantic transfer rule in the cross-domain knowledge transfer channel. The semantic transfer rule is generated in step S135, which stipulates how to transform the source discipline semantic feature into the target discipline semantic feature. According to the semantic transfer rule, each feature element in the source discipline semantic feature is transformed. For example, if the semantic transfer rule stipulates that a certain concept in the source discipline is to be transformed into the corresponding concept in the target discipline, then the feature element representing this concept in the source discipline semantic feature is replaced or adjusted accordingly. During the transformation process, it is necessary to ensure that the generated target semantic feature is aligned with the dimension of the target discipline semantic feature. Dimension alignment can ensure that the target semantic feature can be integrated with the knowledge system of the target discipline. Through the above transformation processing, the source discipline semantic feature is transformed into a target semantic feature that matches the target discipline semantic feature.

[0092] Step S144: Process the source discipline structure features through the structure transfer rules in the cross - domain knowledge transfer channel to generate target structure features that are dimension - aligned with the target discipline structure features.

[0093] Similarly, process the source discipline structure features by using the structure transfer rules in the cross - domain knowledge transfer channel. The structure transfer rules specify how to convert the source discipline structure features into target discipline structure features. According to the structure transfer rules, each feature element in the source discipline structure features is converted. For example, if the structure transfer rule stipulates that the chapter - level structure of the source discipline is to be converted into the code - module hierarchical structure of the target discipline, then the feature elements representing the chapter - level structure in the source discipline structure features are adjusted according to the rule. During the conversion process, it is necessary to ensure that the generated target structure features are dimension - aligned with the target discipline structure features, so that the target structure features can adapt to the knowledge structure of the target discipline. Through the above - mentioned conversion process, the source discipline structure features are converted into target structure features that match the target discipline structure features.

[0094] Step S145: Perform feature fusion processing on the target semantic features and the target structure features to generate a target discipline domain feature representation.

[0095] After obtaining the target semantic features and the target structure features, perform feature fusion processing on them. Feature fusion can integrate the target semantic features and the target structure features together to form a unified target discipline domain feature representation. A splicing method similar to that in Step S125 can be adopted to splice the target semantic feature vector and the target structure feature vector. Since the target semantic features and the target structure features have been dimension - aligned with the semantic features and the structure features of the target discipline respectively, the spliced target discipline domain feature representation can accurately reflect the feature information of the target discipline.

[0096] Step S146: Based on the target discipline domain feature representation, perform an adaptive adjustment process on the target knowledge content of the source discipline literature unit to generate a knowledge transfer result that conforms to the knowledge expression norms of the target discipline. The knowledge transfer result includes the adjusted semantic content and the adjusted structural organization method.

[0097] Based on the generated target discipline domain feature representation, perform an adaptive adjustment process on the target knowledge content of the source discipline literature unit.

[0098] Step S1461: Perform content decomposition processing on the target knowledge content of the source discipline literature unit to separate the core argument part and the auxiliary argument part.

[0099] First, perform content decomposition on the target knowledge content of the source discipline literature unit. Separate the core argument part and the auxiliary argument part from the target knowledge content. The core argument part is the core of the knowledge content, which expresses the main viewpoints and conclusions of the knowledge; the auxiliary argument part is the content such as evidence and examples used to support the core argument. For example, in the target knowledge content of a medical literature, the core argument may be the effectiveness of a treatment method for a certain disease, and the auxiliary argument part may be relevant case analyses, experimental data, etc. The target knowledge content can be decomposed through text analysis techniques according to the importance and logical relationships of sentences.

[0100] Step S1462: According to the target semantic features in the target discipline field feature representation, perform term substitution on the core argument part, replace the source discipline-specific terms with equivalent terms in the target discipline, and generate core argument content with semantic adaptation.

[0101] Perform term substitution on the core argument part according to the target semantic features in the target discipline field feature representation. The source discipline and the target discipline may use different terms to express the same or similar concepts. For example, in the interdisciplinary knowledge transfer between medicine and computer science, the term "disease diagnosis" in medicine may be expressed as "intelligent diagnosis algorithm" in computer science. Through the information in the target semantic features, find the equivalent terms of the source discipline-specific terms in the target discipline, and then replace the source discipline-specific terms in the core argument part with the equivalent terms in the target discipline. This can make the core argument content semantically adapted to the target discipline and facilitate the understanding of readers in the target discipline.

[0102] Step S1463: According to the target structural features in the target discipline field feature representation, perform structural reorganization on the auxiliary argument part, reorganize the presentation order of the auxiliary argument content according to the chapter hierarchy relationship and citation association relationship of the target discipline literature unit, and generate auxiliary argument content with structural adaptation.

[0103] Perform structural reorganization on the auxiliary argument part according to the target structural features in the target discipline field feature representation. The target discipline literature unit may have its specific chapter hierarchy relationship and citation association relationship. In order to make the auxiliary argument content structurally adapted to the target discipline, it is necessary to reorganize the presentation order of the auxiliary argument content according to the structural requirements of the target discipline. For example, the target discipline literature unit may require presenting experimental results first and then analyzing the results, while the auxiliary argument content of the source discipline literature unit may be analyzing first and then presenting the results. According to the target structural features, adjust the order of the auxiliary argument content to generate auxiliary argument content with structural adaptation.

[0104] Step S1464: Perform content coherence check processing on the core argument content of the semantic adaptation and the auxiliary argument content of the structural adaptation to ensure the logical continuity of the two parts of the content.

[0105] Perform content coherence check processing on the core argument content of the semantic adaptation and the auxiliary argument content of the structural adaptation. Although the core argument content and the auxiliary argument content are respectively adapted to the target discipline in terms of semantics and structure after term replacement and structural reorganization, the logical coherence between the two parts of the content may be affected. For example, after term replacement, the logical relationship between the core argument and the auxiliary argument may no longer be clear. Through content coherence check processing, check the logical relationship between the two parts of the content to ensure their logical continuity. A logical reasoning algorithm can be used to analyze whether the logical connection between the two parts of the content is reasonable and adjust the incoherent parts.

[0106] Step S1465: Perform format standardization processing on the content that passes the coherence check, unify the punctuation usage norms and paragraph separation rules, and generate a knowledge transfer result that conforms to the knowledge expression norms of the target discipline.

[0107] After the coherence check passes, perform format standardization processing on the content. Different disciplines may have different punctuation usage norms and paragraph separation rules. For example, in the medical discipline, the use of punctuation may pay more attention to expressing rigorous logical relationships, while in the literature discipline, punctuation may be more used to create an emotional atmosphere. Paragraph separation rules also vary by discipline. Some disciplines may prefer long paragraphs for discussion, while others are more accustomed to short paragraphs.

[0108] Regarding the use of punctuation, it needs to be unified according to the norms of the target discipline. A punctuation usage rule library for the target discipline can be established, and the common punctuation usage methods in the target discipline are stored in it. Then, check each sentence of the core argument content of the semantic adaptation and the auxiliary argument content of the structural adaptation, and replace the non-standard punctuation according to the rules in the rule library. For example, if the target discipline is accustomed to using semicolons to separate parallel sentence elements, while commas are used in the source discipline, then replace the commas with semicolons.

[0109] The unification of paragraph separation rules is also crucial. The paragraph separation standard can be determined according to the paragraph characteristics of the target discipline literature unit. For example, the target discipline literature unit may separate paragraphs at the beginning of each new view or discussion, then adjust the content according to this standard. By analyzing the theme and logical relationship of the sentences, judge whether paragraph separation is needed, and re-segment the content that does not conform to the paragraph separation rules of the target discipline.

[0110] After the unified processing of punctuation usage norms and paragraph separation rules, the generated content is the result of knowledge transfer that conforms to the knowledge expression norms of the target discipline. This knowledge transfer result includes the adjusted semantic content and the adjusted structural organization method, which is not only semantically adapted to the target discipline but also follows the expression habits of the target discipline in terms of structure and format.

[0111] Step S150: Perform a validity verification process on the knowledge transfer result to generate transfer validity evaluation information indicating the reliability of the knowledge transfer.

[0112] To ensure the quality and reliability of knowledge transfer, it is necessary to perform a validity verification process on the knowledge transfer result and generate corresponding transfer validity evaluation information.

[0113] Step S151: Select target discipline benchmark literature units related to the research topic of the knowledge transfer result from the multi-domain academic literature collection. The target discipline benchmark literature units contain authoritative knowledge content verified by domain experts.

[0114] Select target discipline benchmark literature units related to the research topic of the knowledge transfer result from the multi-domain academic literature collection. These benchmark literature units contain authoritative knowledge content verified by domain experts, which can be used as a reference standard for verifying the validity of the knowledge transfer result. For example, in the cross-disciplinary knowledge transfer scenario between medicine and computer science, if the knowledge transfer result is about the intelligent diagnosis algorithm knowledge in the field of computer science, then select authoritative literature units related to the research topic of the intelligent diagnosis algorithm, reviewed and recognized by domain experts, from the computer science literature units as the target discipline benchmark literature units. The appropriate target discipline benchmark literature units can be screened by analyzing information such as the title, abstract, and keywords of the literature units, combined with the recommendations of domain experts.

[0115] Step S152: Extract the domain feature representation of the target discipline benchmark literature unit as the benchmark feature representation, and the benchmark feature representation includes benchmark semantic features and benchmark structural features.

[0116] After selecting the target discipline benchmark literature units, extract their domain feature representations as the benchmark feature representations. This benchmark feature representation includes benchmark semantic features and benchmark structural features. In the previous steps, the domain feature extraction process has been performed on each literature unit in the multi-domain academic literature collection, so the domain feature representation of the target discipline benchmark literature unit can be directly extracted from the stored domain feature representations. The benchmark semantic features reflect the core content semantic information of the target discipline benchmark literature unit, and the benchmark structural features reflect the structural association information of the target discipline benchmark literature unit.

[0117] Step S153: Extract the domain feature representation of the knowledge transfer result as the transfer feature representation, where the transfer feature representation includes transfer semantic features and transfer structural features.

[0118] Similarly, extract the domain feature representation of the knowledge transfer result as the transfer feature representation. This transfer feature representation includes transfer semantic features and transfer structural features. Perform the same domain feature extraction process on the knowledge transfer result as on the previous literature unit, including steps such as text preprocessing, semantic feature extraction, structural feature extraction, and normalization processing. The finally obtained transfer feature representation can reflect the feature information of the knowledge transfer result and is used to compare with the benchmark feature representation.

[0119] Step S154: Calculate the semantic matching degree between the transfer semantic features and the benchmark semantic features, where the semantic matching degree is generated by comparing the overlapping ratio of the feature sub-vectors representing the same research concept in the two features.

[0120] Calculate the semantic matching degree between the transfer semantic features and the benchmark semantic features. The semantic matching degree is used to measure the similarity between the semantic content of the knowledge transfer result and the semantic content of the benchmark literature unit of the target discipline. To calculate the semantic matching degree, first, it is necessary to find the feature sub-vectors representing the same research concept in the transfer semantic features and the benchmark semantic features. A vector matching algorithm, such as the cosine similarity algorithm, can be used to compare each feature sub-vector of the transfer semantic features and the benchmark semantic features pairwise to find the feature sub-vector pairs with higher similarity. Then, count the number of these feature sub-vector pairs representing the same research concept and calculate their overlapping ratio in the total number of feature sub-vectors of the transfer semantic features and the benchmark semantic features. This overlapping ratio is the semantic matching degree, and the closer it is to 1, the more similar the semantic content of the transfer semantic features and the benchmark semantic features.

[0121] Step S155: Calculate the structural matching degree between the transfer structural features and the benchmark structural features, where the structural matching degree is generated by comparing the overlapping ratio of the feature sub-vectors representing the same logical organization method in the two features.

[0122] Calculate the structural matching degree between the migrated structural features and the benchmark structural features. The structural matching degree is used to measure the similarity between the structural organization of the knowledge migration result and the structural organization of the target discipline benchmark literature unit. Similar to calculating the semantic matching degree, first use the vector matching algorithm to find the feature sub-vectors in the migrated structural features and the benchmark structural features that represent the same logical organization method. For example, if both the migrated structural features and the benchmark structural features have feature sub-vectors representing the chapter hierarchical relationship, by comparing the similarity of these feature sub-vectors, find the pairs of feature sub-vectors representing the same chapter hierarchical relationship. Then, count the number of these feature sub-vector pairs representing the same logical organization method, and calculate the overlapping ratio of them in the total number of feature sub-vectors of the migrated structural features and the benchmark structural features. This overlapping ratio is the structural matching degree, and the closer it is to 1, the more similar the structural organization methods of the migrated structural features and the benchmark structural features are.

[0123] Step S156: Calculate the comprehensive matching degree value according to the semantic matching degree and the structural matching degree. The comprehensive matching degree value integrates the semantic matching degree and the structural matching degree in a weighted average manner.

[0124] According to the calculated semantic matching degree and structural matching degree, calculate the comprehensive matching degree value. The comprehensive matching degree value can comprehensively reflect the matching degree of the knowledge migration result with the target discipline benchmark literature unit in terms of both semantics and structure. To calculate the comprehensive matching degree value, weights need to be assigned to the semantic matching degree and the structural matching degree. The assignment of weights needs to consider the importance of semantic content and structural organization in the target discipline. For example, if semantic content is more critical in the target discipline, a higher weight can be assigned to the semantic matching degree; if structural organization is more important for the expression and understanding of knowledge, a higher weight can be assigned to the structural matching degree. Multiply the semantic matching degree by its weight and the structural matching degree by its weight and then add them together to obtain the comprehensive matching degree value.

[0125] Step S157: Generate migration effectiveness evaluation information indicating the reliability of knowledge migration according to the comprehensive matching degree value. The migration effectiveness evaluation information includes the comprehensive matching degree value and the corresponding reliability level description.

[0126] Step S1571: Establish a mapping relationship table between the comprehensive matching degree value and the reliability level. The mapping relationship table includes the reliability level labels corresponding to different comprehensive matching degree value intervals.

[0127] In this step, in order to accurately evaluate the reliability of knowledge transfer based on the comprehensive matching degree value, a mapping relationship table between the comprehensive matching degree value and the reliability level needs to be established. This mapping relationship table associates different intervals of comprehensive matching degree values with corresponding reliability level labels. For example, setting a relatively high interval of comprehensive matching degree values corresponds to the "high reliability" level label, which means that the comprehensive matching degree values within this interval indicate that the knowledge transfer result is very similar to the target discipline benchmark literature unit, and the reliability of knowledge transfer is relatively high; while a relatively low interval of comprehensive matching degree values corresponds to the "low reliability" level label, indicating that the comprehensive matching degree values within this interval represent that the knowledge transfer result has a large difference from the target discipline benchmark literature unit, and the reliability of knowledge transfer is relatively low. The establishment of this mapping relationship table needs to be based on a large amount of experimental data and domain experience to ensure its rationality and accuracy.

[0128] Step S1572: Search the mapping relationship table according to the comprehensive matching degree value to determine the corresponding reliability level label.

[0129] After the mapping relationship table is established, according to the calculated comprehensive matching degree value, search in the mapping relationship table to determine the corresponding reliability level label. By comparing the comprehensive matching degree value with different intervals in the mapping relationship table, find the interval where this comprehensive matching degree value is located, and then determine the corresponding reliability level label. This process can intuitively reflect the reliability of knowledge transfer.

[0130] Step S1573: Extract the semantically mismatched parts in the knowledge transfer result where the semantic matching degree is lower than the preset threshold, and record the specific content of the semantically mismatched parts and the conversion differences between the corresponding source discipline terms and target discipline terms.

[0131] To more comprehensively evaluate the knowledge transfer result, it is necessary to extract the semantically mismatched parts in the knowledge transfer result where the semantic matching degree is lower than the preset threshold. The preset threshold is a pre-set standard used to judge whether the semantic matching degree meets the standard. By comparing the transferred semantic features with the benchmark semantic features, find the parts where the semantic matching degree is lower than this threshold. For these semantically mismatched parts, record their specific content in detail, including relevant text paragraphs, sentences, etc. At the same time, record the conversion differences between the corresponding source discipline terms and target discipline terms, and analyze the possible problems in the term replacement process.

[0132] Step S1574: Extract the structurally mismatched parts in the knowledge transfer result where the structure matching degree is lower than the preset threshold, and record the specific content of the structurally mismatched parts and the organizational differences between the corresponding source discipline structure and target discipline structure.

[0133] Similarly, extract the structurally mismatched parts in the knowledge transfer results where the structural matching degree is lower than the preset threshold. Compare the transferred structural features with the benchmark structural features to identify the parts where the structural matching degree fails to reach the preset threshold. Record the specific content of these structurally mismatched parts, such as the relevant chapter organization, paragraph arrangement, etc. At the same time, analyze the organizational differences between the corresponding source discipline structure and the target discipline structure to understand the possible unreasonable points in the process of structural reorganization, so as to improve the knowledge transfer process.

[0134] Step S1575: Integrate the comprehensive matching degree value, the reliability level label, the record of semantic mismatched parts, and the record of structurally mismatched parts to generate transfer effectiveness evaluation information including quantitative matching data and qualitative difference analysis.

[0135] Finally, integrate the comprehensive matching degree value, the reliability level label, the record of semantic mismatched parts, and the record of structurally mismatched parts. The comprehensive matching degree value and the reliability level label provide quantitative information about the reliability of knowledge transfer, while the record of semantic mismatched parts and the record of structurally mismatched parts provide qualitative difference analysis. Integrating these pieces of information together generates comprehensive transfer effectiveness evaluation information. This evaluation information enables users to clearly understand the effect of knowledge transfer and the problems existing in terms of semantics and structure.

[0136] Throughout the process, it is necessary to strictly follow the principles of legality, fairness, and privacy protection in data processing to ensure the feasibility and compliance of the technical solution. When collecting privacy-sensitive data, technical means such as data encryption and anonymization processing are adopted to prevent the leakage and abuse of privacy-sensitive data. For example, during the text preprocessing of literature units, the parts that may contain privacy-sensitive information are specially marked and processed, protecting the privacy of the data without affecting the effect of knowledge transfer. In the construction and training process of the artificial intelligence model, the necessary modules, layers, and connection relationships of the semantic encoding model, as well as the specific steps and parameters of training, are defined. The pre-trained semantic encoding model includes modules such as a word embedding layer, a bidirectional long short-term memory network layer, an attention mechanism layer, and a fully connected layer. Through training with a large amount of text data, it learns the semantics and context information of natural language. When applying this semantic encoding model, the text fragment unit is used as the input, and the semantic features reflecting the core content of the literature unit are output, realizing the effective combination of the semantic encoding model and the academic literature field.

[0137] Figure 2FIG. 0 shows a schematic diagram of exemplary hardware and software components of a cross - domain knowledge transfer system 100 for academic literature that can implement the ideas of the present application. For example, the processor 120 can be used in the cross - domain knowledge transfer system 100 for academic literature and is used to execute the functions in the present application.

[0138] The cross - domain knowledge transfer system 100 for academic literature can be a general - purpose server or a special - purpose server, both of which can be used to implement the cross - domain knowledge transfer method for academic literature of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0139] For example, the cross - domain knowledge transfer system 100 for academic literature can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the cross - domain knowledge transfer system 100 for academic literature can also include program instructions stored in ROM, RAM, or other types of non - transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The cross - domain knowledge transfer system 100 for academic literature also includes an input / output I / O interface 150 between the computer and other input / output devices.

[0140] For ease of explanation, only one processor is described in the cross - domain knowledge transfer system 100 for academic literature. However, it should be noted that the cross - domain knowledge transfer system 100 in the present application can also include multiple processors. Therefore, the steps executed by one processor described in the present application can also be jointly executed or separately executed by multiple processors. For example, if the processor of the cross - domain knowledge transfer system 100 for academic literature executes step A and step B, it should be understood that step A and step B can also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.

[0141] In addition, an embodiment of the present invention also provides a readable storage medium, in which computer - executable instructions are preset. When the processor executes the computer - executable instructions, the above - mentioned cross - domain knowledge transfer method for academic literature is implemented.

[0142] It should be noted that, in order to simplify the description of the present invention disclosure and thus assist in the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes a plurality of features are merged into one embodiment, drawing or description thereof.

Claims

1. A cross-domain knowledge transfer method for academic literature, characterized in that The method includes: Obtaining a multi-domain academic literature collection containing multiple subject categories, where the multi-domain academic literature collection is composed of literature units corresponding to different subjects, and each literature unit contains text content and domain identification information of the subject to which it belongs; Performing domain feature extraction processing on the multi-domain academic literature collection to obtain a domain feature representation of each literature unit, where the domain feature representation includes semantic features reflecting the core content of the literature unit and structural features reflecting the structural association of the literature unit; Constructing a cross-domain knowledge transfer channel based on the domain feature representation, where the cross-domain knowledge transfer channel is used to establish an association mapping relationship between domain feature representations of different subject fields; Using the cross-domain knowledge transfer channel to transfer the knowledge of the source subject literature unit to the target subject literature unit to generate a knowledge transfer result of the target subject literature unit; Performing validity verification processing on the knowledge transfer result to generate transfer validity evaluation information indicating the reliability degree of knowledge transfer.

2. The cross-domain knowledge transfer method for academic literature according to claim 1, wherein The performing domain feature extraction processing on the multi-domain academic literature collection to obtain a domain feature representation of each literature unit, where the domain feature representation includes semantic features reflecting the core content of the literature unit and structural features reflecting the structural association of the literature unit, includes: Performing text preprocessing operations on each literature unit in the multi-domain academic literature collection, where the text preprocessing operations include removing non-content symbols, filtering common stop words, and splitting continuous text paragraphs into text fragment units that can be processed; Invoking a pre-trained semantic encoding model to perform context semantic analysis processing on the text fragment units to generate semantic features reflecting the core content of the literature unit, where the semantic features include context association information of key concepts in the text fragment units and semantic expression vectors of core arguments; Performing parsing processing on the text structure information of the literature unit, where the text structure information includes the hierarchical relationship of chapter titles, the citation association relationship between charts and the text, and the subject distribution relationship of reference documents; Extracting structural features reflecting the structural association of the literature unit based on the text structure information, where the structural features include the hierarchical depth parameter of chapter titles, the citation frequency parameter between charts and the text, and the subject concentration parameter of reference documents; Performing normalization processing on the semantic features and the structural features to generate a domain feature representation with a unified representation form, where the domain feature representation includes both semantic association information and structural association information.

3. The cross-domain knowledge transfer method for academic literature according to claim 2, wherein The invoking a pre-trained semantic encoding model to perform context semantic analysis processing on the text fragment units to generate semantic features reflecting the core content of the literature unit, includes: Inputting the text fragment units into the word embedding layer of the semantic encoding model to generate an initial word vector representation of the vocabulary in each text fragment unit; Performing temporal context modeling processing on the initial word vector representation through the bidirectional long short-term memory network layer of the semantic encoding model to capture the semantic dependency relationship of the vocabulary in the text fragment units in the context of the previous and subsequent texts, and generating an intermediate semantic vector containing context information; Use the attention mechanism layer of the semantic encoding model to perform key information focusing processing on the intermediate semantic vector, identify the key concept words strongly related to the core content of the literature unit in the text fragment unit, and generate the attention weight distribution of the key concept words; Perform weighted aggregation processing on the intermediate semantic vector according to the attention weight distribution to generate local semantic features reflecting the context association information of the key concept words; Input the local semantic features into the fully connected layer of the semantic encoding model for dimension compression processing to generate a semantic expression vector with a fixed dimension representation, and the semantic expression vector maintains semantic information consistency with the local semantic features.

4. The cross-domain knowledge transfer method for academic literature according to claim 2, wherein The structure features reflecting the structural association of the literature unit are extracted based on the text structure information, and the structure features include the hierarchical depth parameter of the chapter title, the citation frequency parameter of the chart and the text, and the subject concentration parameter of the reference literature, including: Perform hierarchical depth calculation processing on the chapter title hierarchical relationship in the text structure information, count the hierarchical distance of each chapter title relative to the root directory, and generate the hierarchical depth parameter of the chapter title, and the hierarchical depth parameter is used to represent the core degree of the chapter content in the literature unit; Perform citation frequency statistics processing on the citation association relationship between the chart and the text in the text structure information, count the number of times each chart is cited by the text paragraph, and generate the citation frequency parameter of the chart and the text, and the citation frequency parameter is used to represent the closeness of the association between the chart content and the text content; Perform subject concentration calculation processing on the subject distribution relationship of the reference literature in the text structure information, count the proportion of the number of literature belonging to the same subject in the reference literature, and generate the subject concentration parameter of the reference literature, and the subject concentration parameter is used to represent the subject focus degree of the research content of the literature unit; Perform standardization processing on the hierarchical depth parameter, the citation frequency parameter and the subject concentration parameter to eliminate the dimensional difference between different parameters, and generate a set of comparable structure features; Perform feature fusion processing on the set of structure features, and integrate the hierarchical depth parameter, the citation frequency parameter and the subject concentration parameter into a single-dimensional structure feature by weighted summation, and the structure feature retains the original semantic information of each parameter.

5. The cross-domain knowledge transfer method for academic literature according to claim 1, wherein The cross-domain knowledge transfer channel is constructed based on the domain feature representation, and the cross-domain knowledge transfer channel is used to establish the association mapping relationship between the domain feature representations of different subject fields, including: Select cross-disciplinary literature pairs containing common research topics from the multi-domain academic literature collection, and the cross-disciplinary literature pairs are composed of literature units belonging to different disciplines but with overlapping research topics; Extract the domain feature representation of the source subject literature unit and the domain feature representation of the target subject literature unit in the cross-disciplinary literature pair; Calculate the feature similarity value between the domain feature representation of the source subject literature unit and the domain feature representation of the target subject literature unit, and the feature similarity value is calculated by the cosine similarity algorithm; Screen the interdisciplinary literature pairs with similarity values higher than the preset threshold according to the described feature similarity values as key mapping samples; Perform association rule mining on the source discipline domain feature representation and the target discipline domain feature representation of the key mapping samples to generate a set of transfer rules reflecting the cross-disciplinary feature correspondence; Perform structured storage processing on the set of transfer rules to generate a cross-domain knowledge transfer channel for establishing an association mapping relationship between feature representations in different discipline domains.

6. The cross-domain knowledge transfer method for academic literature according to claim 5, wherein The performing association rule mining on the source discipline domain feature representation and the target discipline domain feature representation of the key mapping samples to generate a set of transfer rules reflecting the cross-disciplinary feature correspondence includes: Perform feature decomposition on the source discipline domain feature representation of the key mapping samples to separate the semantic feature part and the structural feature part therein; Perform synchronous feature decomposition on the target discipline domain feature representation of the key mapping samples to separate the corresponding semantic feature part and the structural feature part; Establish a semantic correspondence between the source discipline semantic feature part and the target discipline semantic feature part, and the semantic correspondence is generated by matching the feature sub-vectors representing the same research concept in the two parts; Establish a structural correspondence between the source discipline structural feature part and the target discipline structural feature part, and the structural correspondence is generated by matching the feature sub-vectors representing the same logical organization method in the two parts; Perform rule induction on the semantic correspondence and the structural correspondence to generate a set of transfer rules including semantic transfer rules and structural transfer rules. The semantic transfer rules are used to guide the conversion of source discipline semantic features to target discipline semantic features, and the structural transfer rules are used to guide the conversion of source discipline structural features to target discipline structural features; Perform conflict detection on the set of transfer rules to eliminate contradictory descriptions between different rules and generate the final set of transfer rules.

7. The cross-domain knowledge transfer method for academic literature according to claim 1, characterized in that The using the cross-domain knowledge transfer channel to transfer the knowledge of the source discipline literature unit to the target discipline literature unit to generate the knowledge transfer result of the target discipline literature unit includes: Select the source discipline literature unit and the target discipline literature unit that need to perform knowledge transfer from the multi-domain academic literature collection, and the source discipline literature unit contains the target knowledge content to be transferred; Extract the domain feature representation of the source discipline literature unit, and the domain feature representation includes the source discipline semantic feature and the source discipline structural feature; Perform conversion processing on the source discipline semantic feature through the semantic transfer rules in the cross-domain knowledge transfer channel to generate a target semantic feature aligned with the target discipline semantic feature dimension; Perform conversion processing on the source discipline structural feature through the structural transfer rules in the cross-domain knowledge transfer channel to generate a target structural feature aligned with the target discipline structural feature dimension; Perform feature fusion on the target semantic feature and the target structural feature to generate a target discipline domain feature representation; Based on the representation of the target subject field characteristics, perform an adaptive adjustment process on the target knowledge content of the source subject literature unit to generate a knowledge transfer result that conforms to the target subject knowledge expression specification. The knowledge transfer result includes the adjusted semantic content and the adjusted structural organization method.

8. The cross-domain knowledge transfer method for academic literature according to claim 7, wherein The process of performing an adaptive adjustment process on the target knowledge content of the source subject literature unit based on the representation of the target subject field characteristics to generate a knowledge transfer result that conforms to the target subject knowledge expression specification includes: Perform a content decomposition process on the target knowledge content of the source subject literature unit to separate the core argument part and the auxiliary argument part therein; According to the target semantic characteristics in the representation of the target subject field characteristics, perform a term replacement process on the core argument part, replacing the source subject-specific terms with target subject equivalent terms to generate a semantically adapted core argument content; According to the target structural characteristics in the representation of the target subject field characteristics, perform a structural reorganization process on the auxiliary argument part, reorganizing the presentation order of the auxiliary argument content according to the chapter hierarchy relationship and citation association relationship of the target subject literature unit to generate a structurally adapted auxiliary argument content; Perform a content coherence check process on the semantically adapted core argument content and the structurally adapted auxiliary argument content, and perform a format standardization process on the content that passes the coherence check, unifying the punctuation usage specification and paragraph separation rules to generate a knowledge transfer result that conforms to the target subject knowledge expression specification.

9. The cross-domain knowledge transfer method for academic literature according to claim 1, wherein The process of performing an effectiveness verification process on the knowledge transfer result to generate transfer effectiveness evaluation information indicating the reliability degree of the knowledge transfer includes: Select a target subject benchmark literature unit related to the research topic of the knowledge transfer result from the multi-field academic literature collection. The target subject benchmark literature unit includes authoritative knowledge content verified by domain experts; Extract the field characteristic representation of the target subject benchmark literature unit as the benchmark characteristic representation. The benchmark characteristic representation includes benchmark semantic characteristics and benchmark structural characteristics; Extract the field characteristic representation of the knowledge transfer result as the transfer characteristic representation. The transfer characteristic representation includes transfer semantic characteristics and transfer structural characteristics; Calculate the semantic matching degree between the transfer semantic characteristics and the benchmark semantic characteristics. The semantic matching degree is generated by comparing the overlapping ratio of the feature sub-vectors representing the same research concept in the transfer semantic characteristics and the benchmark semantic characteristics; Calculate the structural matching degree between the transfer structural characteristics and the benchmark structural characteristics. The structural matching degree is generated by comparing the overlapping ratio of the feature sub-vectors representing the same logical organization method in the transfer structural characteristics and the benchmark structural characteristics; Calculate the comprehensive matching degree value according to the semantic matching degree and the structural matching degree. The comprehensive matching degree value integrates the semantic matching degree and the structural matching degree by means of weighted average; Generate transfer effectiveness evaluation information indicating the reliability degree of the knowledge transfer according to the comprehensive matching degree value. The transfer effectiveness evaluation information includes the comprehensive matching degree value and the corresponding reliability level description.

10. A cross - domain knowledge transfer system for academic literature, characterized in that, It includes a processor and a memory. The memory is connected to the processor. The memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the cross-domain knowledge transfer method for academic literature according to any one of claims 1-9 above.

Citation Information

Patent Citations

  • Scientific and technological literature content depth revealing method based on content map

    CN111651562A

  • Care question and answer model training method based on transfer learning and expert feedback

    CN117542471A

  • Domain speech recognition method and system based on RAG

    CN119296516A

  • Subject cross measurement method based on keywords

    CN119443903A

  • Acquisition and application of contextual role knowledge for coreference resolution

    US20090326919A1

Cited By

  • Cross-domain knowledge migration method and device, electronic equipment, medium and product

    CN120872326A