Semantic association analysis method and system based on academic literature

By acquiring and analyzing the basic semantic features of academic literature, and using pre-trained models to generate semantic correlation feature representations between documents, it solves the problem of difficulty in digging deep into semantic correlation between documents in the prior art, and realizes efficient literature correlation analysis and retrieval.

CN120337936AActive Publication Date: 2025-07-18JIEHELIX (SHANGHAI) MEDICAL TECH CO LTD

Patent Information

Application Number
CN202510816613.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

Existing academic literature analysis methods are difficult to deeply explore the semantic relationships and logical relationships between documents, resulting in inefficiency and susceptibility to subjective factors.

Method used

By obtaining the data set of academic literature texts, basic semantic features are extracted, pre-trained semantic correlation analysis model is called for correlation feature modeling, semantic correlation feature representations between documents are generated, and output to the target literature management system to support association retrieval.

Benefits of technology

It realizes in-depth mining and accurate analysis of complex semantic relationships between academic literature, improves automation level and accuracy, helps scientific researchers quickly discover potential connections, and promotes academic innovation and knowledge dissemination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337936A_ABST
    Figure CN120337936A_ABST
Patent Text Reader

Abstract

The invention provides a semantic association analysis method and system based on academic literatures, and aims to solve the problem that semantic association between literatures is difficult to deeply mine in an existing academic literature analysis method.The method comprises the steps that firstly, an academic literature text data set to be analyzed is obtained, and basic semantic feature extraction is conducted; core concept features and context semantic relation features of each document are obtained; then, calling a pre-trained semantic association analysis model to carry out association feature modeling on the basic semantic features, and generating semantic association feature representations among the literatures; then, determining semantic association types and association strength among the literatures according to the semantic association feature representation; finally, analysis feedback information containing the literature association relationship is generated and output to the target literature management system to support literature content association retrieval, automatic and accurate analysis of semantic association between academic literatures is achieved, and improvement of the efficiency and quality of academic research is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a semantic association analysis method and system based on academic literature. Background Art

[0002] In the field of academic research, with the rapid development of information technology and the increasing abundance of digital literature resources, the number of academic literature has shown an explosive growth. These literatures cover various subject fields and provide valuable knowledge sources for scientific researchers. However, in the face of a vast amount of academic literature, how to efficiently mine the internal connections between literatures, discover potential research trends and knowledge associations, has become an important challenge faced by current academic research.

[0003] Traditional academic literature analysis methods mainly rely on manual reading and expert judgment. This method is not only inefficient but also easily affected by subjective factors, making it difficult to comprehensively and accurately reveal the semantic associations between literatures. Although natural language processing technologies and machine learning algorithms have been applied to a certain extent in the field of academic literature analysis in recent years, most of the existing methods focus on the superficial analysis of literature content, such as keyword extraction, text classification, etc., and fail to deeply mine the semantic associations and logical relationships between literatures. Summary of the Invention

[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, embodiments of the present invention provide a semantic association analysis method based on academic literature, and the method includes: Obtain a set of academic literature text data to be analyzed, where the set of academic literature text data includes multiple literature text units with topic tags; Perform basic semantic feature extraction processing on the set of academic literature text data to obtain basic semantic features of each literature text unit, where the basic semantic features include core concept features and context semantic relationship features of the literature content; Call a pre-trained semantic association analysis model to perform association feature modeling processing on the basic semantic features to generate a semantic association feature representation between literature text units, where the semantic association feature representation includes the concept co-occurrence association degree and semantic logical connection degree of the literature content; Determine the semantic association analysis result in the set of academic literature text data according to the semantic association feature representation, where the semantic association analysis result includes the description information of the semantic association type and association strength between literatures; Generate analysis feedback information including literature association relationships based on the semantic association analysis result, and output the analysis feedback information to a target literature management system to support the literature content association retrieval operation.

[0005] In another aspect, an embodiment of the present invention further provides a semantic association analysis system based on academic literature, including a processor and a machine-readable storage medium. The machine-readable storage medium is connected to the processor. The machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.

[0006] Based on the above aspects, the embodiment of the present invention realizes the in-depth mining and accurate analysis of the complex semantic associations between academic documents by comprehensively applying technical means such as basic semantic feature extraction, association feature modeling, and semantic association analysis. It can not only automatically extract the core concept features and context semantic relationship features of the document content, but also effectively capture the concept co-occurrence association degree and semantic logic connection degree between documents through a pre-trained semantic association analysis model, and generate an accurate semantic association feature representation. Based on the semantic association feature representation, it is possible to further determine the semantic association type and the description information of the association strength between documents, providing intuitive and comprehensive document association analysis results for scientific research personnel, significantly improving the automation level and accuracy of academic document semantic association analysis, and helping scientific research personnel quickly discover the potential connections between documents, grasp research trends, and promote academic innovation and knowledge dissemination. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 is a schematic flowchart of the execution of the semantic association analysis method based on academic literature provided by an embodiment of the present invention.

[0008] Figure 2 is a schematic diagram of an exemplary hardware and software component of the semantic association analysis system based on academic literature provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0009] The present invention will be specifically described below with reference to the accompanying drawings of the specification. Figure 1 is a schematic flowchart of the semantic association analysis method based on academic literature provided by an embodiment of the present invention. The semantic association analysis method based on academic literature will be introduced in detail below.

[0010] Step S110: Obtain a set of academic document text data to be analyzed, where the set of academic document text data includes multiple document text units with topic tags.

[0011] In this embodiment, the academic literature text data set consists of multiple literature text units, and each literature text unit is attached with a topic tag. The topic tag is a general identifier for the core content of the literature. For example, in the field of medical research, the topic tags can be "treatment of cardiovascular diseases", "tumor immunotherapy", etc. Among them, these literature text units can be obtained through various channels, such as academic databases, professional library websites, etc. From these data sources, the literature text units that meet the requirements are screened according to the set search conditions. For example, in a well-known academic database, keywords and topic tags are used for searching, and the literature text units with the topic tag of "treatment of cardiovascular diseases" found are collected to form an academic literature text data set. These literature text units may come from different eras, different research institutions and authors, covering various research perspectives and achievements under this topic.

[0012] Step S120: Perform basic semantic feature extraction processing on the academic literature text data set to obtain the basic semantic features of each literature text unit. The basic semantic features include the core concept features and context semantic relationship features of the literature content.

[0013] After obtaining the academic literature text data set, the next step is to perform basic semantic feature extraction processing on it to obtain the basic semantic features of each literature text unit. The basic semantic features mainly consist of core concept features and context semantic relationship features. The core concept features represent the most critical concepts and ideas in the literature, while the context semantic relationship features reflect the logical relationships and semantic connections of these concepts in the literature context. For example, in a literature about "the application of artificial intelligence in medical image diagnosis", the core concept features may include "artificial intelligence algorithms", "medical images", "diagnostic models", etc., and the context semantic relationship features reflect the interactions between these concepts, such as how "artificial intelligence algorithms" are applied to the processing of "medical images" and then a "diagnostic model" is constructed.

[0014] Step S121: Perform text cleaning processing on the academic literature text data set to remove the format mark information and non-content symbol information in the literature text units, and obtain pure text data units with unified formats.

[0015] When performing basic semantic feature extraction, first, text cleaning processing needs to be performed on the academic literature text data set. The literature text units may contain some format mark information and non-content symbol information during storage and transmission, which will interfere with subsequent semantic analysis, so they need to be removed. The format mark information includes header and footer marks, chapter title format marks, formula number marks, etc., and the non-content symbol information includes special symbols, hyperlink addresses, footnote reference marks, etc. The specific operations are as follows: Step S1211: Identify the format tag information in the literature text unit, where the format tag information includes header and footer tags, chapter title format tags, and formula number tags.

[0016] In this step, it is necessary to identify the format tag information in the literature text unit. Different literatures may adopt different format tagging methods, but commonly there are specific symbol combinations or typesetting rules to represent headers and footers, chapter titles, and formula numbers. For example, headers and footers may appear at the top and bottom of the page in a specific font, font size, or position; chapter titles may use specific title level tags, such as "1. Introduction", "2. Related Work", etc.; formula numbers may appear in the form of parentheses and numbers next to the formula, such as "(1)", "(2)", etc. These format tag information can be identified by methods such as regular expression matching and text pattern recognition. Regular expressions are a powerful text matching tool that can find and identify specific content in text according to a preset pattern. For example, for chapter title format tags, the regular expression "\d+\.[\u4e00-\u9fa5a-zA-Z]+" can be used to match chapter titles that start with "", followed by a number and a dot, and then Chinese or English characters.

[0017] Step S1212: Perform batch deletion processing on the format tag information through regular expression matching, and retain the body content information in the literature text unit.

[0018] After identifying the format tag information, use regular expression matching to perform batch deletion processing on it. Regular expressions can accurately find the format tag information according to a preset pattern and replace it with an empty string or other specified characters. For example, for the chapter title format tags identified earlier, after using regular expression matching, replace them with an empty string, so that the format tags of the chapter titles can be removed, and only the body content information is retained. During the processing, ensure that the pattern of the regular expression is set accurately to avoid accidentally deleting the body content. At the same time, for some complex format tags, different regular expressions may need to be used multiple times for matching and deletion to ensure that all format tag information is removed.

[0019] Step S1213: Identify the non-content symbol information in the literature text unit, where the non-content symbol information includes special symbols, hyperlink addresses, and footnote reference tags.

[0020] In addition to format marker information, the document text unit also contains some non-content symbol information, such as special symbols, hyperlink addresses, and footnote reference markers, etc. Special symbols include various punctuation marks, mathematical symbols, emojis, etc. Hyperlink addresses usually start with "http: / / " or "https: / / ", and footnote reference markers may appear in the text in the form of numbers or other symbols. These non-content symbol information can be identified by methods such as character matching and string searching. For example, for hyperlink addresses, they can be identified by searching for strings starting with "http: / / " or "https: / / "; for footnote reference markers, they can be identified by searching for specific combinations of numbers or symbols in the text.

[0021] Step S1214: Perform character-by-character detection processing on the non-content symbol information, and replace the detected non-content symbol information with blank characters.

[0022] After identifying the non-content symbol information, perform character-by-character detection processing on it. Character-by-character detection can ensure accurate identification and processing of each non-content symbol. For the detected non-content symbol information, replace it with blank characters. For example, for special symbols, they can be replaced with spaces; for hyperlink addresses, the whole can be replaced with a single space; for footnote reference markers, they can also be replaced with spaces. During the replacement process, pay attention to maintaining the coherence of the text to avoid incomplete or semantically chaotic text due to the replacement operation.

[0023] Step S1215: Perform blank line merging processing on the processed document text unit, merge consecutive multiple blank lines into a single blank line, and obtain a pure text data unit with a unified format.

[0024] After the previous processing, there may be consecutive multiple blank lines in the document text unit. To make the text format more unified, perform blank line merging processing on the processed document text unit. It can be achieved by traversing each line of the text, detecting consecutive blank lines, and merging them into a single blank line. For example, when detecting that two consecutive lines are both blank lines, delete one of them and only keep one blank line. Through the above method, a pure text data unit with a unified format is finally obtained, providing clean and standardized text data for subsequent semantic analysis.

[0025] Step S122: Perform word segmentation processing on the pure text data unit, extract the core term unit and key phrase unit in the document text unit, where the core term unit contains professional terms in the document research field, and the key phrase unit contains descriptive phrases of the logical relationship between terms.

[0026] After obtaining the plain text data units with unified formats, perform word segmentation on them to extract the core term units and key phrase units in the literature text units. Word segmentation is to split continuous text into individual words or phrases according to set rules. Core term units are professional terms in the literature research field, which represent the core concepts and research directions of the literature; key phrase units contain descriptive phrases of the logical relationships between terms, which reflect the semantic connections and logical structures between core terms. For example, in a literature about "gene editing technology", the core term units may include "CRISPR / Cas9", "gene editing", "off-target effect", etc., and the key phrase units may include "CRISPR / Cas9-based gene editing", "methods for reducing off-target effect", etc. Professional word segmentation tools can be used, such as Jieba for Chinese and NLTK for English, to perform word segmentation on the plain text data units. During the word segmentation process, according to the research field and professional characteristics of the literature, screen and adjust the word segmentation results to ensure accurate extraction of core term units and key phrase units. For example, for some professional terms in specific fields, they may need to be added to the dictionary of the word segmentation tool to improve the accuracy of word segmentation.

[0027] Step S123: Invoke the pre-trained semantic encoding model to perform context semantic encoding processing on the core term units and key phrase units, generating an initial semantic vector of the literature text unit, where the initial semantic vector contains the co-occurrence relationship between term units and the logical cohesion relationship of phrase units.

[0028] After extracting the core term units and key phrase units, invoke the pre-trained semantic encoding model to perform context semantic encoding processing on them to generate an initial semantic vector of the literature text unit. The pre-trained semantic encoding model is trained through a large amount of text data and can learn the semantic relationships between words and phrases. The initial semantic vector contains the co-occurrence relationship between term units and the logical cohesion relationship of phrase units, and these logical cohesion relationships reflect the semantic information of the literature text unit. The specific operations are as follows: Step S1231: Arrange the core term units and key phrase units in the order of appearance in the literature text unit to generate a term phrase sequence.

[0029] First, arrange the extracted core term units and key phrase units in the order of their appearance in the literature text unit to generate a term phrase sequence. Such an arrangement can retain the context order of terms and phrases in the original text, which helps with subsequent semantic encoding processing. For example, in a literature about "quantum computing", after arranging the core term units and key phrase units in the order of appearance, the generated term phrase sequence may be "qubit", "quantum algorithm", "design of quantum algorithm based on qubit", etc.

[0030] Step S1232: Input the term phrase sequence into the word embedding layer of the semantic encoding model to generate corresponding word vector representations for each term unit and phrase unit, where the word vector representations contain the basic semantic information of the term phrases.

[0031] Input the generated term phrase sequence into the word embedding layer of the semantic encoding model. The role of the word embedding layer is to convert each term unit and phrase unit into corresponding word vector representations. A word vector is a numerical vector that can map the semantic information of words and phrases into a low-dimensional vector space. Each word vector representation contains the basic semantic information of the term phrase, and the positions and distances of different word vectors in the vector space reflect their semantic similarities. For example, in the semantic encoding model, the word vectors of "qubit" and "quantum bit" may be close in the vector space because they have similar semantics.

[0032] Step S1233: Perform sequence modeling processing on the word vector representations through the context encoder of the semantic encoding model to capture the context dependencies of the term phrases in the sequence and generate context encoding vectors containing context semantic information.

[0033] After obtaining the word vector representations, perform sequence modeling processing on them through the context encoder of the semantic encoding model. The context encoder can capture the context dependencies of the term phrases in the sequence, that is, consider the context information of the terms and phrases in the original text, so as to generate context encoding vectors containing context semantic information. For example, in a sentence "Quantum computing utilizes the characteristics of qubits for high-speed computing", the context encoding vector of "qubit" will consider the context information such as "quantum computing" and "high-speed computing", so as to more accurately reflect its semantics in the sentence. The context encoder can adopt model structures such as recurrent neural network (RNN), long short-term memory network (LSTM), gated recurrent unit (GRU), etc. These models can effectively process sequence data and capture the long-term dependencies in the sequence.

[0034] Step S1234: Perform attention weighting processing on the context encoding vectors, assign attention weights according to the importance of the term phrases in the literature text unit, enhance the semantic representations of the key term phrases, and weaken the semantic representations of the secondary term phrases.

[0035] To highlight the semantic information of key term phrases, attention weighting is performed on the generated context encoding vectors. The attention mechanism can assign attention weights according to the importance of term phrases in the literature text units. Important term phrases will receive higher attention weights, while less important term phrases will receive lower attention weights. Through attention weighting, the semantic representation of key term phrases can be enhanced, and the semantic representation of less important term phrases can be weakened. For example, in a literature on "optimization of artificial intelligence algorithms", key terms such as "genetic algorithm" and "particle swarm algorithm" may receive higher attention weights, while some auxiliary descriptive phrases may receive lower attention weights. The attention mechanism can be implemented in various ways, such as dot product attention, additive attention, etc.

[0036] Step S1235: Perform dimension concatenation processing on the attention-weighted context encoding vectors to generate the initial semantic vector of the literature text unit.

[0037] After performing attention weighting, perform dimension concatenation processing on the attention-weighted context encoding vectors. Dimension concatenation is to connect multiple vectors together according to a set rule to form a higher-dimensional vector. Through dimension concatenation processing, the context encoding vectors of each term phrase are integrated together to generate the initial semantic vector of the literature text unit. The initial semantic vector contains the co-occurrence relationship between term units and the logical connection relationship of phrase units, and can comprehensively reflect the semantic information of the literature text unit.

[0038] Step S124: Perform feature screening processing on the initial semantic vector, retain the semantic dimension information strongly related to the literature theme label, and remove the redundant semantic dimension information weakly related to the literature theme label to obtain the core concept features of the literature text unit.

[0039] After generating the initial semantic vector of the literature text unit, perform feature screening processing on it. The initial semantic vector may contain some redundant semantic dimension information weakly related to the literature theme label. These information will increase the complexity of subsequent analysis and may interfere with the extraction of core concepts. Therefore, it is necessary to retain the semantic dimension information strongly related to the literature theme label and remove the redundant information to obtain the core concept features of the literature text unit. Feature screening can be performed by calculating the correlation between each dimension of the initial semantic vector and the literature theme label. Correlation calculation can use methods such as cosine similarity and Pearson correlation coefficient. For example, for a literature on "new energy vehicles", its theme label is "technological innovation of new energy vehicles". The dimensions related to "new energy" and "automobile technology" in the initial semantic vector may have a higher correlation with the theme label, while the dimensions in other unrelated fields have a lower correlation. Remove the dimensions with lower correlation and retain the dimensions with higher correlation to obtain the core concept features.

[0040] Step S125: Analyze the positional distribution relationship of the core term units and key phrase units in the literature text units, and extract the context semantic relationship features of the literature content. The context semantic relationship features include the proximity relationship of term units within a paragraph and the progressive relationship of phrase units between chapters.

[0041] In addition to the core concept features, it is also necessary to extract the context semantic relationship features of the literature content. The context semantic relationship features can reflect the positional distribution relationship of term units and key phrase units in the literature text units, including the proximity relationship of term units within a paragraph and the progressive relationship of phrase units between chapters. For example, in a paragraph, adjacent term units may have a closer semantic connection; in different chapters, the order of appearance of key phrase units may reflect the progressive relationship of the research content. The context semantic relationship features can be extracted by analyzing the positional information of the core term units and key phrase units in the literature text units, such as the paragraphs, sentences, and chapters where they are located. For the proximity relationship of term units within a paragraph, the distance between term units can be calculated, and the closer the distance, the stronger the proximity relationship; for the progressive relationship of phrase units between chapters, the order of appearance and logical connection of phrase units in different chapters can be analyzed.

[0042] Step S126: Standardize the core concept features and the context semantic relationship features to generate basic semantic features with a unified dimensional representation.

[0043] After obtaining the core concept features and context semantic relationship features, they need to be standardized. After standardization, the core concept features and context semantic relationship features are concatenated to generate basic semantic features with a unified dimensional representation. The basic semantic features contain the core concept features and context semantic relationship features of the literature content, and can comprehensively reflect the semantic information of the literature text units.

[0044] Step S130: Invoke a pre-trained semantic association analysis model to perform association feature modeling on the basic semantic features, and generate a semantic association feature representation between the literature text units. The semantic association feature representation includes the concept co-occurrence association degree and semantic logical connection degree of the literature content.

[0045] After obtaining the basic semantic features of each document text unit, a pre-trained semantic association analysis model is called to perform association feature modeling on it. The semantic association analysis model is pre-trained with a large amount of document data and can learn the semantic association patterns between documents. By processing the basic semantic features through this model, a semantic association feature representation between document text units is generated, which includes the concept co-occurrence association degree and semantic logical connection degree of the document content. The concept co-occurrence association degree reflects the co-occurrence of core concepts in different documents, and the semantic logical connection degree reflects the semantic logical relationship between different documents. For example, in a group of documents about "big data analysis", if both documents mention "data mining algorithms", their concept co-occurrence association degree is relatively high; if the research conclusion of one document is the basis for the research of another document, their semantic logical connection degree is relatively high.

[0046] Step S131: Input the basic semantic features into the feature input layer of the semantic association analysis model for feature standardization processing to obtain standardized basic semantic features.

[0047] First, input the obtained basic semantic features into the feature input layer of the semantic association analysis model. The main function of the feature input layer is to preprocess the input basic semantic features, including feature standardization processing. Feature standardization processing can unify the basic semantic features of different document text units to the same scale, avoiding problems such as unstable model training or inaccurate results caused by different feature scales. Multiple methods can be used for feature standardization, such as z-score standardization and min-max standardization. Z-score standardization converts the feature values to a standard normal distribution by calculating the mean and standard deviation of the features; min-max standardization scales the feature values to a fixed interval, such as [0, 1]. Through feature standardization processing, standardized basic semantic features are obtained, providing unified input data for subsequent association feature modeling.

[0048] Step S132: Through the local association modeling layer of the semantic association analysis model, perform in-document local association analysis on the standardized basic semantic features to extract the local association patterns of the core concept features and the context semantic relationship features in each document text unit, where the local association pattern includes the importance weights of the concept features in the context relationship.

[0049] After obtaining the standardized basic semantic features, perform in-document local association analysis on them through the local association modeling layer of the semantic association analysis model. The purpose of the local association modeling layer is to extract the local association patterns of the core concept features and the context semantic relationship features in each document text unit. The specific operations are as follows: Step S1321: For the standardized basic semantic features of each document text unit, separate the core concept features and the context semantic relationship features.

[0050] The standardized basic semantic features obtained after feature standardization processing contain these two parts of information: core concept features and context semantic relationship features. In order to perform subsequent local association analysis, these two parts of features need to be separated. Since the representation method or positional relationship of the core concept features and the context semantic relationship features in the standardized basic semantic features has been clarified in the previous steps, they can be separated based on this information. For example, if the standardized basic semantic feature is in the form of a vector, and it is known that the first part of the dimensions represents the core concept features and the second part of the dimensions represents the context semantic relationship features, then the two parts of features can be extracted respectively according to this dimension division rule.

[0051] Step S1322: Calculate the dot product similarity between each concept dimension in the core concept features and each relationship dimension in the context semantic relationship features to generate a similarity matrix.

[0052] After separating the core concept features and the context semantic relationship features, it is necessary to calculate the similarity between them. Here, the calculation method of dot product similarity is adopted. For each concept dimension in the core concept features, a dot product operation is performed with each relationship dimension in the context semantic relationship features. The dot product operation can measure the similarity between two vectors. In this case, it is to measure the semantic similarity between the concept dimension and the relationship dimension. By performing dot product operations on all concept dimensions and relationship dimensions, the obtained results are arranged into a similarity matrix. For example, if the core concept features have m concept dimensions and the context semantic relationship features have n relationship dimensions, then the generated similarity matrix is an m-row and n-column matrix, and each element in the matrix represents the dot product similarity between a concept dimension and a relationship dimension.

[0053] Step S1323: Perform row normalization on the similarity matrix to obtain the importance weights of each concept dimension in different relationship dimensions.

[0054] After obtaining the similarity matrix, in order to clarify the importance of each concept dimension in different relationship dimensions, it is necessary to perform row normalization on the similarity matrix. Row normalization is to process each row element of the matrix so that the sum of each row element is 1. Through row normalization processing, the similarity of each concept dimension in different relationship dimensions can be transformed into importance weights. Because after normalization, the value of each row element represents the relative importance degree of the concept dimension in each relationship dimension. For example, for the row corresponding to a certain concept dimension, the larger the value of a certain element, the higher the importance of the concept dimension in the corresponding relationship dimension.

[0055] Step S1324: Perform column normalization on the similarity matrix to obtain the coverage range weights of each relationship dimension in different concept dimensions.

[0056] In addition to calculating the importance weights of each concept dimension in different relationship dimensions, it is also necessary to calculate the coverage range weights of each relationship dimension in different concept dimensions. This requires performing column normalization on the similarity matrix. Column normalization is to process each column element of the matrix so that the sum of each column element is 1. Through column normalization, the similarity of each relationship dimension in different concept dimensions can be transformed into coverage range weights. After normalization, the value of each column element represents the relative coverage degree of the relationship dimension in each concept dimension. For example, for a column corresponding to a certain relationship dimension, the larger the value of a certain element, the wider the coverage range of the relationship dimension in the corresponding concept dimension.

[0057] Step S1325: Based on the importance weights and the coverage range weights, construct a local association pattern of the core concept features and the context semantic relationship features in the literature text unit. The local association pattern includes the relationship importance distribution of the concept dimension and the concept coverage distribution of the relationship dimension.

[0058] After obtaining the importance weights of each concept dimension in different relationship dimensions and the coverage range weights of each relationship dimension in different concept dimensions, a local association pattern of the core concept features and the context semantic relationship features in the literature text unit can be constructed. This local association pattern combines the relationship importance distribution of the concept dimension and the concept coverage distribution of the relationship dimension. The relationship importance distribution of the concept dimension describes the importance degree of each concept dimension in different context relationships, and the concept coverage distribution of the relationship dimension describes the coverage range of each context relationship for different concept dimensions. Through this local association pattern, the local association situation between the core concept features and the context semantic relationship features within the literature can be clearly understood, such as which concepts are more important in certain contexts and which context relationships cover more concepts.

[0059] Step S133: Perform global association analysis processing on the standardized basic semantic features through the global association modeling layer of the semantic association analysis model, calculate the similarity matching degree of the core concept features and the cohesion consistency of the context semantic relationship features between different literature text units, and obtain the global association matching result between the literatures.

[0060] After completing the local association analysis within the literature, it is necessary to further analyze the global association situation between text units of different literatures. The global association modeling layer of the semantic association analysis model is responsible for performing global association analysis processing on the standardized basic semantic features across literatures. For the core concept features between text units of different literatures, calculate their similarity matching degrees. Multiple methods can be used to calculate the similarity matching degrees, such as cosine similarity, Euclidean distance, etc. Cosine similarity can measure the cosine value of the angle between two vectors. The closer the value is to 1, the more similar the two vectors are, which means the core concept features of the two literatures are more similar. Euclidean distance calculates the spatial distance between two vectors, and the smaller the distance, the more similar the core concept features of the two literatures are. For the context semantic relationship features, calculate their coherence consistency. This can be achieved by analyzing aspects such as the logical order and semantic coherence of the context relationships in different literatures. For example, if the context relationship of one literature unfolds sequentially according to the research steps, and another literature follows a similar logical order when discussing the same topic, then it can be considered that the context semantic relationship features of these two literatures have a high coherence consistency. By calculating the similarity matching degrees of the core concept features and the coherence consistency of the context semantic relationship features, the global association matching result between literatures is finally obtained.

[0061] Step S134: Input the local association pattern and the global association matching result into the feature fusion layer of the semantic association analysis model, and based on the attention mechanism, dynamically allocate the fusion weights of the local association pattern and the global association matching result to generate a fused association feature vector.

[0062] After obtaining the local association pattern within the literature and the global association matching result between literatures, it is necessary to fuse these two parts of information. The feature fusion layer of the semantic association analysis model is responsible for completing this fusion task. The attention mechanism is used to dynamically allocate the fusion weights of the local association pattern and the global association matching result. The attention mechanism can automatically adjust the importance degrees of the local association pattern and the global association matching result in the fusion process according to different situations. For example, in some cases, the local association information within the literature may be more important, then the attention mechanism will assign a higher weight to the local association pattern; in other cases, the global association information between literatures may be more critical, and the attention mechanism will assign a higher weight to the global association matching result. By dynamically allocating the fusion weights, the local association pattern and the global association matching result are weighted and concatenated to generate a fused association feature vector. This fused association feature vector integrates the local association information within the literature and the global association information between literatures.

[0063] Step S135: Perform feature dimensionality reduction on the fused association feature vector, retain the key dimensionality information that can represent the semantic association between documents, and generate a semantic association feature representation between document text units. The semantic association feature representation includes the concept co-occurrence association degree and the semantic logical connection degree.

[0064] The generated fused association feature vector may have a relatively high dimensionality, which contains some redundant information that has little effect on representing the semantic association between documents. In order to simplify the data and highlight the key information, it is necessary to perform feature dimensionality reduction on the fused association feature vector. There are various methods for feature dimensionality reduction, such as principal component analysis (PCA), linear discriminant analysis (LDA), etc. Principal component analysis is to project high-dimensional data into a low-dimensional space by finding the principal components of the data while retaining the variance information of the data as much as possible; linear discriminant analysis is to perform dimensionality reduction by finding the projection direction that can maximize the separation between different categories. When performing feature dimensionality reduction, retain the key dimensionality information that can represent the semantic association between documents. After the dimensionality reduction process, a semantic association feature representation between document text units is generated. This semantic association feature representation includes the concept co-occurrence association degree and the semantic logical connection degree. The concept co-occurrence association degree reflects the degree of co-occurrence of core concepts in different documents, and the semantic logical connection degree reflects the coherence degree of semantic logic between different documents.

[0065] Step S140: Determine the semantic association analysis result in the academic document text data set according to the semantic association feature representation. The semantic association analysis result includes the semantic association type between documents and the description information of the association strength.

[0066] After obtaining the semantic association feature representation between document text units, it is necessary to further determine the semantic association analysis result in the academic document text data set. The semantic association analysis result includes the semantic association type between documents and the description information of the association strength.

[0067] Step S141: Analyze the concept co-occurrence association degree and the semantic logical connection degree in the semantic association feature representation, and extract the distribution information of the association degree values between document text units.

[0068] First, analyze the semantic association feature representation to extract the information of the concept co-occurrence association degree and the semantic logical connection degree. The concept co-occurrence association degree and the semantic logical connection degree exist in the semantic association feature representation in the form of numerical values. By extracting these numerical information and organizing them into the distribution information of the association degree values. The distribution information of the association degree values describes the distribution of the association degrees between different document text units, such as which documents have a higher association degree and which documents have a lower association degree, etc.

[0069] Step S142: Perform clustering analysis on the correlation degree numerical distribution information to identify groups of literature text units with strong correlation and groups of literature text units with weak correlation.

[0070] The purpose of this step is to distinguish groups of literature text units with strong and weak correlations based on the correlation degree numerical distribution information through clustering analysis. The following are the specific sub-steps and detailed descriptions: Step S1421: Select the density clustering algorithm as the clustering analysis method and set the clustering radius parameter and the minimum sample number parameter.

[0071] Among various clustering algorithms, the density clustering algorithm can perform clustering based on the density distribution of data points, which is suitable for processing data with different density regions and is more applicable to data such as the correlation degree numerical distribution information that may have an irregular distribution. The clustering radius parameter is used to determine the neighborhood range of each data point, and the size of this clustering radius parameter will affect the number of data points within the neighborhood; the minimum sample number parameter stipulates the minimum number of neighborhood data points required for a data point to become a core point. The settings of these two parameters need to be adjusted according to the distribution characteristics of the correlation degree values. If the correlation degree values are relatively dispersed, a larger clustering radius and a smaller minimum sample number may be required; conversely, a smaller clustering radius and a larger minimum sample number are needed. For example, in a research field where the correlation degree values between literature vary greatly, these two parameters need to be adjusted flexibly to obtain accurate clustering results.

[0072] Step S1422: Use the correlation degree value of each pair of literature text units in the correlation degree numerical distribution information as a sample point to construct a sample point set.

[0073] The correlation degree value of each pair of literature text units is the basic data unit for clustering analysis. Regarding these correlation degree values as sample points and integrating all the sample points together form a sample point set. This sample point set is the basis for subsequent clustering operations and contains information on the correlation degrees between all pairs of literature. For example, in an academic literature text data set containing multiple medical literatures, there is a correlation degree value between every two literatures. Collecting these values to construct a sample point set covers the correlation information of all pairs of literatures in the entire literature set.

[0074] Step S1423: Calculate the Euclidean distance between each sample point in the sample point set and other sample points to determine the neighborhood sample points of each sample point.

[0075] Euclidean distance is a commonly used distance metric that can measure the distance between two sample points in space. For each sample point in the set of sample points, calculate its Euclidean distance from all other sample points. According to the pre-set clustering radius parameter, determine the neighborhood range of each sample point, and the sample points within the neighborhood are the neighborhood sample points of that sample point. By calculating the Euclidean distance and determining the neighborhood sample points, the data distribution around each sample point can be understood. For example, in a set of sample points composed of the correlation values of different subject literatures, after calculating the Euclidean distance between each sample point and other sample points, it can be known which sample points are closer and may belong to the same clustering group.

[0076] Step S1424: Determine the type of the sample point according to the number of the neighborhood sample points, and the types of the sample points include core points, boundary points, and noise points.

[0077] According to the number of neighborhood sample points of the sample point, the sample point can be divided into core points, boundary points, and noise points. If the number of neighborhood sample points of a sample point is greater than or equal to the minimum sample number parameter, then the sample point is a core point. The core point is the center of the clustering and represents a data area with a higher density. If the number of neighborhood sample points of a sample point is less than the minimum sample number parameter, but it is within the neighborhood of a certain core point, then the sample point is a boundary point, and the boundary points surround the core points. If a sample point is neither a core point nor a boundary point, then it is a noise point. Noise points are usually isolated data points and do not belong to any clustering group. For example, in a set of sample points about the correlation of scientific and technological literatures, some sample points have many neighborhood sample points around them, and these are core points; some sample points have fewer neighborhood sample points around them but are close to core points, and these are boundary points; while those isolated sample points are noise points.

[0078] Step S1425: Merge the core point and its neighborhood sample points into a clustering group, assign the boundary points to the clustering group to which the nearest core point belongs, and exclude the noise points.

[0079] After determining the type of the sample point, merge the core point and its neighborhood sample points into a clustering group. The core point is the center of the clustering, and its neighborhood sample points have a strong correlation with the core point, so they are grouped together. For boundary points, since they are close to a certain core point, they are assigned to the clustering group to which the nearest core point belongs, which can make the clustering result more complete. And since the noise points have a weak correlation with other data points and have little impact on the clustering result, they are excluded. For example, in a set of sample points containing the correlation of different research directions in computer science, merge each core point and its neighborhood sample points and the corresponding boundary points into different clustering groups respectively. After excluding the noise points, a clear clustering result is obtained.

[0080] Step S1426: According to the average value of the correlation degree values of the sample points in the clustering group, mark the clustering group with an average correlation degree value higher than the preset threshold as the literature text unit group with strong correlation degree, and mark the clustering group with an average correlation degree value lower than the preset threshold as the literature text unit group with weak correlation degree.

[0081] The preset threshold is the boundary for distinguishing strong correlation and weak correlation. Calculate the average value of the correlation degree values of the sample points in each clustering group, and compare this average value with the preset threshold. If the average value is higher than the preset threshold, it indicates that the correlation degree among the literature text units in this clustering group is relatively high, and mark it as the literature text unit group with strong correlation degree; conversely, if the average value is lower than the preset threshold, then mark it as the literature text unit group with weak correlation degree. The setting of the preset threshold needs to comprehensively consider the research purpose and data characteristics. For example, in an association analysis of biomedical research literature, if you want to more strictly screen out strongly associated literature groups, you can appropriately increase the preset threshold.

[0082] Step S143: Analyze the distribution characteristics of the concept co-occurrence correlation degree and semantic logical connection degree in the literature text unit group with strong correlation degree, and determine the semantic association type between the literatures. The semantic association type includes the concept extension association type and the logical argument association type.

[0083] After identifying the literature text unit group with strong correlation degree, further analyze the distribution characteristics of the concept co-occurrence correlation degree and semantic logical connection degree in this group. For the concept co-occurrence correlation degree, if a large number of the same core concepts frequently appear in certain literatures, and these concepts are further expanded and deepened in different literatures, then it can be determined that there is a concept extension association type between these literatures. For example, in a series of literatures on "artificial intelligence natural language processing", core concepts such as "word vector model" and "attention mechanism" frequently appear, and different literatures conduct research and expansion on these concepts from different angles, then these literatures belong to the concept extension association type. For the semantic logical connection degree, if there is an obvious logical derivation and argument relationship between certain literatures, for example, one literature proposes a theoretical hypothesis, and another literature verifies this hypothesis through experiments, then it can be determined that there is a logical argument association type between these literatures. By detailed analysis of the concept co-occurrence correlation degree and semantic logical connection degree in the literature text unit group with strong correlation degree, the semantic association type between the literatures can be accurately determined.

[0084] Step S144: Calculate the average value and variance value of the correlation degree values of each pair of literature text units in the literature text unit group with strong correlation degree, and generate association strength description information. The association strength description information includes the correlation degree central tendency parameter and the dispersion degree parameter.

[0085] To more comprehensively describe the association strength among documents in a group of literature text units with a strong association degree, calculate the average value and variance value of the association degree values for each pair of literature text units. The average value can reflect the central tendency of the association degree, that is, the overall level of the association degree among documents in this group; the variance value can reflect the dispersion degree of the association degree, that is, the fluctuation of the association degree among documents in this group. For example, if the average value is high and the variance value is small, it indicates that the association degree among most documents in this group is relatively high and relatively stable; if the average value is high but the variance value is large, it indicates that there are significant differences in the association degree among documents in this group. Generate association strength description information through the average value and variance value, and the association strength description information includes the central tendency parameter of the association degree (i.e., the average value) and the dispersion degree parameter (i.e., the variance value).

[0086] Step S145: Combine the semantic association type and the association strength description information to generate the semantic association analysis result in the academic literature text data set.

[0087] Combine the determined semantic association type and the generated association strength description information to generate the semantic association analysis result in the academic literature text data set. The semantic association analysis result synthesizes the semantic association types among documents (such as the concept extension association type, the logical argumentation association type) and the association strength description information (such as the central tendency parameter of the association degree and the dispersion degree parameter). Through this semantic association analysis result, the semantic association situation among different documents in the academic literature text data set can be clearly understood, including which association type they belong to and the strength of the association.

[0088] Step S150: Generate analysis feedback information containing document association relationships based on the semantic association analysis result, and output the analysis feedback information to the target document management system to support the document content association retrieval operation.

[0089] After obtaining the semantic association analysis result, it is necessary to generate analysis feedback information containing document association relationships based on this result.

[0090] Step S151: Extract the semantic association type and the association strength description information from the semantic association analysis result, and construct structured description data of document association relationships. The structured description data includes association document pair identifiers and corresponding association type labels.

[0091] Extract the semantic association type and the description information of the association strength from the semantic association analysis results. Based on this information, construct the structured description data of the literature association relationship. The structured description data records the association relationship between literatures in a standardized manner, which includes the identifier of the associated literature pair and the corresponding association type label. The identifier of the associated literature pair is used to uniquely identify each pair of literatures with an association relationship, and the association type label specifies the semantic association type between this pair of literatures, such as the concept extension association type, the logical argument association type, etc. By constructing the structured description data, the complex literature association relationship can be represented in a clear and orderly manner.

[0092] Step S152: Perform visual transformation processing on the structured description data to convert the literature association relationship into a visual association graph containing nodes and edges. The nodes represent literature text units, and the edges represent the semantic association relationships between literatures. The attributes of the edges include the association type label and the description information of the association strength.

[0093] In this step, the structured literature association relationship data needs to be converted into an intuitive visual association graph. The following are the specific sub-steps and detailed descriptions: For example, step S1521: Create a node for each literature text unit and an edge for each literature association relationship.

[0094] In the visual association graph, the nodes are the intuitive representations of the literature text units, and each node represents a literature. The edges are used to represent the semantic association relationships between literatures, and there is an edge connecting each pair of literatures with an association relationship. By constructing the nodes and edges, the basic framework of the visual association graph is initially built. For example, in an association analysis of literatures in different sub-fields of physics, each literature corresponds to a node, and the association relationship between literatures corresponds to an edge, so that the complex literature association relationship can be presented in a graphical form.

[0095] Step S1522: Configure node attribute information for the nodes. The node attribute information includes the literature title, the subject label, and the publication time information.

[0096] Node attribute information can enrich the information of the literature represented by the node. The literature title allows users to quickly identify the main content of the literature; the subject tags clarify the subject area to which the literature belongs, helping users to grasp the research direction of the literature from a macro perspective; the publication time information can reflect the timeliness of the literature, which is very important for some users who need to pay attention to the latest research results. In the visualization graph, when users interact with the nodes, these attribute information can be displayed. For example, in a visualization association graph of new energy research literature, when a user clicks on a node, they can see the title, subject tags (such as solar energy, wind energy, etc.) and publication time of the literature, so as to better understand the basic situation of the literature.

[0097] Step S1523: Configure edge attribute information for the edge, and the edge attribute information includes an association type tag and an association strength description information.

[0098] The edge attribute information further describes the association relationship between the literatures. The association type tag clarifies the semantic association type between the literatures. For example, the concept extension association type means that one literature expands on a certain concept based on another literature; the logical argument association type means that one literature provides logical support for the view of another literature. The association strength description information reflects the closeness of the association between the literatures and can be represented by means such as an association degree value. By configuring these attribute information for the edge, users can more deeply understand the nature and strength of the association between the literatures. For example, in a visualization association graph of artificial intelligence algorithm research literature, the association type tag of the edge may be displayed as "algorithm improvement association", and the association strength description information can be represented by a value to indicate the closeness of the association between the two literatures in terms of algorithm improvement.

[0099] Step S1524: Perform color coding processing on the nodes according to the subject tags of the literature text units. Nodes with the same subject tag use the same color, and nodes with different subject tags use different colors.

[0100] Color coding is an intuitive visualization means. Through the distinction of colors, users can quickly identify the literatures with the same subject tag. Nodes with the same subject tag use the same color, so in the graph, the literature nodes with the same research subject will gather together in the same color, facilitating users to grasp the distribution of literatures with different subjects as a whole. Nodes with different subject tags use different colors, which can clearly distinguish the differences between each subject. For example, in a visualization association graph covering literatures in multiple disciplinary fields, the literature nodes in the field of computer science can be represented by blue, and the literature nodes in the field of biology can be represented by green, so that users can immediately see the distribution and association of literatures in different disciplines.

[0101] Step S1525: Perform width encoding processing on the edges according to the association strength description information. The higher the association strength, the wider the edge; the lower the association strength, the narrower the edge.

[0102] The width encoding of the edges can visually display the differences in the association strength between documents. Edges with high association strength are wider and more prominent in the graph, indicating a close association between the two documents; edges with low association strength are narrower and relatively less obvious, indicating a weak association between the two documents. In the above way, users can quickly judge the magnitude of the association strength between documents by the width of the edges. For example, in a visual association graph of financial market research literature, the edge between two documents with a close association in market trend prediction will be wider, while the edge between documents with a weak association will be narrower.

[0103] Step S1526: Arrange the nodes and edges in spatial positions according to the force-directed layout algorithm to generate a visual association graph containing nodes and edges.

[0104] The force-directed layout algorithm is based on the mechanical principles in physics, regarding the nodes as charged particles and the edges as springs. There is a repulsive force between nodes, and the edges generate an attractive force. By continuously iterating and adjusting the positions of the nodes, the entire visual association graph reaches a balanced state. In this process, nodes with a close association will approach each other, and nodes with a weak association will move away from each other. The visual association graph generated in this way can visually display the association structure and density relationship between documents. For example, in a visual association graph of literature research literature, through the force-directed layout algorithm, the document nodes of the same literary genre will gather together, and the document nodes of different literary genres will be relatively dispersed, forming a clear document association network.

[0105] Step S153: Configure an interactive operation interface for the visual association graph. The interactive operation interface supports click query operations on nodes and filtering operations on edges.

[0106] To improve the practicality and interactivity of the visual association graph, an interactive operation interface is configured for it. The interactive operation interface supports click query operations on nodes. When the user clicks on a certain node, the detailed information of the document corresponding to the node can be popped up, such as the document title, subject tags, publication time, core content summary, etc., which is convenient for users to further understand the specific situation of the document. The interactive operation interface also supports filtering operations on edges. Users can filter the edges according to conditions such as association type tags and association strength, and only display the association relationships that meet specific conditions. In this way, it is possible to focus on the document association relationships that users are interested in and improve the efficiency of information retrieval.

[0107] Step S154: Perform data encapsulation processing on the structured description data and the visual association map to generate analysis feedback information containing literature association relationships.

[0108] Perform data encapsulation processing on the structured description data and the visual association map. Data encapsulation is to integrate different types of data together to form a unified data packet. Through data encapsulation processing, the structured description data and the visual association map are combined into a complete analysis feedback information containing literature association relationships. This analysis feedback information contains both the structured data of literature association relationships and the intuitive visual map, providing users with multi-dimensional literature association information.

[0109] Step S155: Perform format adaptation processing on the analysis feedback information to make it meet the data input format requirements of the target literature management system.

[0110] After generating the analysis feedback information, it needs to be output to the target literature management system to support the literature content association retrieval operation. Since different literature management systems may have different data input format requirements, it is necessary to perform format adaptation processing on the analysis feedback information. According to the data input format specifications of the target literature management system, corresponding conversion and adjustment are performed on the analysis feedback information, such as changing the data storage format, adjusting the data field order, etc. Through format adaptation processing, it is ensured that the analysis feedback information can be smoothly input into the target literature management system.

[0111] During the whole process, if there is a situation involving the collection of privacy-sensitive data, for example, when obtaining the academic literature text data set, it may contain some privacy-sensitive data such as the author's personal information. Data desensitization technology can be used to perform privacy protection and anti-disclosure processing on these privacy-sensitive data. Data desensitization technology includes methods such as replacement, masking, and encryption. Replacement is to replace sensitive data with virtual values, for example, replacing the author's real name with an anonymous number; masking is to partially hide sensitive data, for example, only showing the last four digits of the author's mobile phone number; encryption is to use an encryption algorithm to encrypt sensitive data, and only authorized users can decrypt and view it. Through these technical means, privacy can be effectively protected and anti-disclosure can be achieved. In addition to data desensitization technology, access control technology can also be used to strictly manage the permissions of personnel who can access these privacy-sensitive data. Only authorized users can access specific levels of privacy-sensitive data. At the same time, during the data transmission process, use an encrypted transmission protocol, such as the SSL / TLS protocol, to ensure that the data is not stolen or tampered with during transmission. For the server storing privacy-sensitive data, use a secure storage system, perform regular data backups and security audits, and promptly discover and handle potential security risks.

[0112] In the construction and training of artificial intelligence models, pre-trained semantic encoding models and semantic association analysis models play a crucial role in the entire technical solution. For the semantic encoding model, its necessary modules include a word embedding layer, a context encoder, and an attention mechanism module. The word embedding layer is responsible for converting the input core term units and key phrase units into corresponding word vector representations, which is the basis for subsequent semantic encoding. The context encoder usually adopts a recurrent neural network structure such as LSTM or GRU to capture the context dependencies of term phrases in the sequence and generate context encoding vectors containing context semantic information. The attention mechanism module then weights the context encoding vectors to highlight the semantic representations of key term phrases.

[0113] The hierarchical structure of the semantic encoding model is as follows: The input layer receives the term phrase sequence generated by arranging the core term units and key phrase units in order; the word embedding layer converts the input term phrases into word vectors; the context encoder performs sequence modeling on the word vectors; the attention mechanism module weights the context encoding vectors; and finally, the attention-weighted context encoding vectors are output, and the initial semantic vectors of the literature text units are obtained through dimension concatenation.

[0114] When training the semantic encoding model, a large amount of academic literature text data can be used as training data, and these data need to be preprocessed in a similar way to the previous steps, including operations such as text cleaning and word segmentation. The specific training steps are as follows: First, divide the preprocessed text data into a training set, a validation set, and a test set. The training set is used for parameter learning of the model, the validation set is used to adjust the hyperparameters of the model, and the test set is used to evaluate the final performance of the model. Set the training parameters, such as the learning rate, batch size, number of training epochs, etc. The learning rate controls the step size of parameter updates in each iteration, the batch size determines the number of data samples used in each training, and the number of training epochs represents the number of times the entire training data is traversed by the model.

[0115] During the training process, a stochastic gradient descent (SGD) or its variant algorithms, such as the Adam optimization algorithm, are used to update the parameters of the model. Each time a batch of data is taken from the training set and input into the model, the output of the model is calculated through forward propagation, and then the gradient is calculated according to the loss function value between the output and the true label through the backpropagation algorithm. Finally, the parameters of the model are updated according to the gradient. This process is continuously repeated until the performance of the model on the validation set reaches stability or meets the preset stopping conditions.

[0116] For the semantic association analysis model, its necessary modules include a feature input layer, a local association modeling layer, a global association modeling layer, a feature fusion layer, and a feature dimensionality reduction layer. The feature input layer normalizes the input basic semantic features; the local association modeling layer extracts the local association patterns of the core concept features and the context semantic relationship features within the literature; the global association modeling layer calculates the global association matching results between different literature text units; the feature fusion layer fuses the local association patterns and the global association matching results based on the attention mechanism; the feature dimensionality reduction layer reduces the dimensionality of the fused feature vectors to generate the semantic association feature representation between the literature text units.

[0117] The hierarchical structure of the semantic association analysis model is as follows: the input layer receives the normalized basic semantic features; the local association modeling layer and the global association modeling layer perform local and global association analyses respectively; the feature fusion layer fuses the results of the two; the feature dimensionality reduction layer reduces the dimensionality of the fused results; and finally, the semantic association feature representation between the literature text units is output.

[0118] When training the semantic association analysis model, the training data is the academic literature text data after basic semantic feature extraction. The training steps are similar to those of the semantic encoding model, and the data also needs to be divided into a training set, a validation set, and a test set. Set the training parameters, such as the learning rate, batch size, number of training epochs, etc. Use an appropriate optimization algorithm to update the parameters of the model. During the training process, the loss function can be designed according to the characteristics of the semantic association feature representation. For example, the loss function can be constructed by combining the differences between the true values and the model prediction values of the concept co-occurrence association degree and the semantic logical connection degree.

[0119] When applying these artificial intelligence models in a specific field or scenario, taking the academic literature management and retrieval scenario as an example, the input data of the model is the academic literature text data after preprocessing and feature extraction, including the basic semantic features of each literature text unit. The output data of the model is the semantic association feature representation between the literature text units, the semantic association analysis results generated based on this, and the analysis feedback information containing the literature association relationships. The internal association relationship between these input and output data is that through the model's processing of the input basic semantic features, the semantic association information between the literatures is mined and finally output in the form of analysis feedback information to provide support for the association retrieval operation of the literature management system.

[0120] In terms of data collection, ensure that the data collection process complies with laws, regulations, and ethical norms. For the collection of academic literature text data, it can be obtained through legal academic database interfaces. When obtaining the data, the authorization of the database owner needs to be obtained, and the usage terms of the database need to be complied with. At the same time, during the collection process, the data is preliminarily screened and filtered to remove the obviously non-compliant literatures, such as duplicate literatures, low-quality literatures, etc.

[0121] In terms of tag management, accurate topic tags are assigned to each document text unit. The assignment of topic tags can adopt a combination of manual annotation and automatic annotation. Manual annotation is carried out by professional domain experts to ensure the accuracy and professionalism of the tags. Automatic annotation can use machine learning algorithms to automatically assign topic tags to documents by extracting features and classifying the document text. At the same time, a tag management system is established to uniformly manage and maintain the topic tags to ensure the consistency and standardization of the tags.

[0122] In terms of rule setting, clear rules are formulated to guide the entire semantic association analysis process. For example, in the feature extraction process, rules for word segmentation are set to determine which words should be extracted as core term units and key phrase units; in the clustering analysis process, the value ranges and adjustment rules of the clustering radius parameter and the minimum sample number parameter are set. The setting of these rules needs to be reasonably adjusted and optimized according to the specific application scenarios and data characteristics.

[0123] In terms of recommendation decision-making, based on the generated semantic association analysis results and analysis feedback information, recommendations for document association retrieval are provided to users. For example, when a user searches for a certain document in the document management system, the system can recommend related documents according to the semantic association relationship between this document and other documents. The rules for recommendation decision-making can be set according to the association type and association strength, and documents with strong association degrees and specific association types (such as concept extension association types) are recommended first.

[0124] During the implementation process of the entire technical solution, performance evaluation and optimization should be continuously carried out. By evaluating the output results of the model, such as calculating indicators such as accuracy, recall rate, and F1 value, the performance of the model is evaluated. According to the evaluation results, the parameters, structure, or training process of the model are adjusted and optimized to improve the performance and accuracy of the model. At the same time, the process of the entire semantic association analysis method is optimized, such as adjusting the feature extraction method and optimizing the clustering analysis algorithm, to improve the efficiency and effect of the entire technical solution.

[0125] In practical applications, further expansion and utilization of the analysis feedback information can also be considered. For example, the analysis feedback information can be combined with the citation relationship of documents to construct a more comprehensive document knowledge graph. The document knowledge graph can intuitively display the semantic association relationship and citation relationship between documents, providing more in-depth knowledge discovery and research support for academic researchers. At the same time, an intelligent retrieval system based on the document knowledge graph can be developed to support more complex document association retrieval operations, such as multi-condition retrieval based on semantic association and citation relationship, and association path query.

[0126] For the storage and management of analysis feedback information, a dedicated database needs to be established to store the structured descriptive data containing literature association relationships and the visualized association graphs. The design of the database should consider data storage efficiency, query efficiency, and data security. Relational databases or non-relational databases, such as MySQL, MongoDB, etc., can be used and selected according to the characteristics of the data and application requirements. In the database, corresponding table structures are established for different types of data to ensure the orderly organization and management of data.

[0127] Figure 2 FIG. shows a schematic diagram of exemplary hardware and software components of a semantic association analysis system 100 based on academic literature that can implement the ideas of the present application provided by some embodiments of the present application. For example, the processor 120 can be used on the semantic association analysis system 100 based on academic literature and is used to execute the functions in the present application.

[0128] The semantic association analysis system 100 based on academic literature can be a general-purpose server or a special-purpose server, both of which can be used to implement the semantic association analysis method based on academic literature of the present application. Although only one server is shown in the present application, for convenience, the functions described in the present application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.

[0129] For example, the semantic association analysis system 100 based on academic literature can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and different forms of storage media 140, such as disks, ROM, or RAM, or any combination thereof. Exemplarily, the semantic association analysis system 100 based on academic literature can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The semantic association analysis system 100 based on academic literature also includes an I / O interface 150 between the computer and other input / output devices.

[0130] For ease of explanation, only one processor is described in the semantic association analysis system 100 based on academic literature. However, it should be noted that the semantic association analysis system 100 in the present application can also include multiple processors. Therefore, the steps executed by one processor described in the present application can also be jointly executed or separately executed by multiple processors. For example, if the processor of the semantic association analysis system 100 based on academic literature executes steps A and B, it should be understood that steps A and B can also be jointly executed by two different processors or separately executed in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor jointly execute steps A and B.

[0131] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the semantic association analysis method based on academic literature as described above is implemented.

[0132] It should be noted that, in order to simplify the description of the present invention disclosure and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, sometimes multiple features are merged into one embodiment, drawing, or description thereof.

Claims

1. A semantic association analysis method based on academic literature, characterized in that The method includes: Obtaining a set of academic literature text data to be analyzed, where the set of academic literature text data contains multiple literature text units with topic tags; Performing basic semantic feature extraction processing on the set of academic literature text data to obtain the basic semantic features of each literature text unit, where the basic semantic features include the core concept features and context semantic relationship features of the literature content; Invoking a pre-trained semantic association analysis model to perform association feature modeling processing on the basic semantic features to generate a semantic association feature representation between literature text units, where the semantic association feature representation includes the concept co-occurrence association degree and semantic logical connection degree of the literature content; Determining the semantic association analysis result in the set of academic literature text data according to the semantic association feature representation, where the semantic association analysis result includes the semantic association type and association strength description information between the literatures; Generating analysis feedback information including literature association relationships based on the semantic association analysis result, and outputting the analysis feedback information to a target literature management system to support literature content association retrieval operations.

2. The semantic association analysis method based on academic literature according to claim 1, characterized in that The performing basic semantic feature extraction processing on the set of academic literature text data to obtain the basic semantic features of each literature text unit includes: Performing text cleaning processing on the set of academic literature text data to remove format marking information and non-content symbol information in the literature text units, and obtaining uniformly formatted pure text data units; Performing word segmentation processing on the pure text data units to extract the core term units and key phrase units in the literature text units, where the core term units include professional terms in the literature research field, and the key phrase units include logical relationship description phrases between the terms; Invoking a pre-trained semantic encoding model to perform context semantic encoding processing on the core term units and key phrase units to generate an initial semantic vector of the literature text unit, where the initial semantic vector includes the co-occurrence relationship between the term units and the logical connection relationship of the phrase units; Performing feature screening processing on the initial semantic vector, retaining the semantic dimension information strongly related to the literature topic tag, and removing the redundant semantic dimension information weakly related to the literature topic tag to obtain the core concept features of the literature text unit; Analyzing the positional distribution relationship of the core term units and key phrase units in the literature text unit, and extracting the context semantic relationship features of the literature content, where the context semantic relationship features include the proximity relationship of the term units within a paragraph and the progressive relationship of the phrase units between chapters; Performing standardization processing on the core concept features and the context semantic relationship features to generate basic semantic features with a unified dimension representation.

3. The semantic association analysis method based on academic literature according to claim 2, wherein The performing text cleaning processing on the set of academic literature text data to remove format marking information and non-content symbol information in the literature text units, and obtaining uniformly formatted pure text data units includes: Identifying the format marking information in the literature text unit, where the format marking information includes header and footer marks, chapter title format marks, and formula number marks; Batch delete the format tag information through a regular expression matching method, and retain the text content information in the literature text unit; Identify the non-content symbol information in the literature text unit, where the non-content symbol information includes special symbols, hyperlink addresses, and footnote reference marks; Perform character-by-character detection processing on the non-content symbol information, and replace the detected non-content symbol information with blank characters; Perform blank line merging processing on the processed literature text unit, merge multiple consecutive blank lines into a single blank line, and obtain a pure text data unit with a unified format.

4. The semantic association analysis method based on academic literature according to claim 2, wherein The pre-trained semantic encoding model is called to perform context semantic encoding processing on the core term unit and the key phrase unit to generate an initial semantic vector of the literature text unit, including: Arrange the core term unit and the key phrase unit in the order of appearance in the literature text unit to generate a term phrase sequence; Input the term phrase sequence into the word embedding layer of the semantic encoding model to generate corresponding word vector representations for each term unit and phrase unit, and the word vector representations contain the basic semantic information of the term phrases; Perform sequence modeling processing on the word vector representations through the context encoder of the semantic encoding model to capture the context dependence relationship of the term phrases in the sequence, and generate context encoding vectors containing context semantic information; Perform attention weighting processing on the context encoding vectors, assign attention weights according to the importance of the term phrases in the literature text unit, enhance the semantic representation of the key term phrases, and weaken the semantic representation of the secondary term phrases; Perform dimension concatenation processing on the attention-weighted context encoding vectors to generate an initial semantic vector of the literature text unit.

5. The semantic association analysis method based on academic literature according to claim 1, characterized in that The pre-trained semantic association analysis model is called to perform association feature modeling processing on the basic semantic features to generate a semantic association feature representation between literature text units, including: Input the basic semantic features into the feature input layer of the semantic association analysis model for feature standardization processing to obtain standardized basic semantic features; Perform local association analysis processing within the literature on the standardized basic semantic features through the local association modeling layer of the semantic association analysis model, extract the local association patterns of the core concept features and the context semantic relationship features in each literature text unit, and the local association patterns include the importance weights of the concept features in the context relationship; Perform global association analysis processing between literatures on the standardized basic semantic features through the global association modeling layer of the semantic association analysis model, calculate the similarity matching degree of the core concept features and the coherence consistency of the context semantic relationship features between different literature text units, and obtain the global association matching result between literatures; Input the local association pattern and the global association matching result into the feature fusion layer of the semantic association analysis model, and dynamically assign fusion weights to the local association pattern and the global association matching result based on the attention mechanism to generate a fused association feature vector; Perform feature dimensionality reduction on the fused correlation feature vectors, retain the key dimensional information that can represent the semantic correlation between documents, and generate the semantic correlation feature representation between document text units. The semantic correlation feature representation includes the concept co-occurrence correlation degree and the semantic logic connection degree.

6. The semantic association analysis method based on academic literature according to claim 5, wherein The local correlation analysis of the standardized basic semantic features is performed by the local correlation modeling layer of the semantic correlation analysis model, and the local correlation patterns of the core concept features and the context semantic relationship features in each document text unit are extracted, including: For the standardized basic semantic features of each document text unit, separate the core concept features and the context semantic relationship features; Calculate the dot product similarity between each concept dimension in the core concept features and each relationship dimension in the context semantic relationship features to generate a similarity matrix; Perform row normalization on the similarity matrix to obtain the importance weights of each concept dimension in different relationship dimensions; Perform column normalization on the similarity matrix to obtain the coverage weights of each relationship dimension in different concept dimensions; Based on the importance weights and the coverage weights, construct the local correlation patterns of the core concept features and the context semantic relationship features in the document text unit. The local correlation patterns include the relationship importance distribution of the concept dimensions and the concept coverage distribution of the relationship dimensions.

7. The semantic association analysis method based on academic literature according to claim 1, characterized in that The semantic correlation analysis results in the academic document text data set are determined according to the semantic correlation feature representation, including: Analyze the concept co-occurrence correlation degree and the semantic logic connection degree in the semantic correlation feature representation, and extract the correlation degree numerical distribution information between document text units; Perform clustering analysis on the correlation degree numerical distribution information to identify the groups of document text units with strong correlation degrees and the groups of document text units with weak correlation degrees; Analyze the distribution characteristics of the concept co-occurrence correlation degree and the semantic logic connection degree in the groups of document text units with strong correlation degrees to determine the semantic correlation types between documents. The semantic correlation types include the concept extension correlation type and the logical argumentation correlation type; Calculate the average value and variance value of the correlation degree numerical values of each pair of document text units in the groups of document text units with strong correlation degrees to generate the correlation strength description information. The correlation strength description information includes the correlation degree central tendency parameter and the dispersion degree parameter; Combine the semantic correlation type and the correlation strength description information to generate the semantic correlation analysis results in the academic document text data set.

8. The semantic association analysis method based on academic literature according to claim 7, characterized in that The clustering analysis is performed on the correlation degree numerical distribution information to identify the groups of document text units with strong correlation degrees and the groups of document text units with weak correlation degrees, including: Select the density clustering algorithm as the clustering analysis method and set the clustering radius parameter and the minimum sample number parameter; Use the correlation degree numerical values of each pair of document text units in the correlation degree numerical distribution information as sample points to construct a sample point set; Calculate the Euclidean distance between each sample point in the sample point set and other sample points to determine the neighborhood sample points of each sample point; Judge the type of the sample point according to the number of neighborhood sample points. The types of the sample points include core points, boundary points, and noise points; Merge the core points and their neighboring sample points into a clustering group, assign the boundary points to the clustering group to which the nearest core point belongs, and exclude the noise points; Based on the average value of the correlation degree values of the sample points in the clustering group, mark the clustering group with an average correlation degree value higher than the preset threshold as a literature text unit group with a strong correlation degree, and mark the clustering group with an average correlation degree value lower than the preset threshold as a literature text unit group with a weak correlation degree.

9. The semantic association analysis method based on academic literature according to claim 1, wherein The analysis feedback information including the literature association relationship generated based on the semantic association analysis result includes: Extract the semantic association type and the association strength description information in the semantic association analysis result, and construct the structured description data of the literature association relationship. The structured description data includes the identifier of the associated literature pair and the corresponding association type label; Perform visual transformation processing on the structured description data, and convert the literature association relationship into a visual association graph including nodes and edges. The nodes represent literature text units, the edges represent the semantic association relationships between the literatures, and the attributes of the edges include the association type label and the association strength description information; Configure an interactive operation interface for the visual association graph. The interactive operation interface supports the click query operation on the nodes and the filtering operation on the edges; Perform data encapsulation processing on the structured description data and the visual association graph to generate the analysis feedback information including the literature association relationship; Perform format adaptation processing on the analysis feedback information to make it meet the data input format requirements of the target literature management system.

10. A semantic association analysis system based on academic literature, characterized in that, It includes a processor and a memory. The memory is connected to the processor. The memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the semantic association analysis method based on academic literature according to any one of claims 1-9 above.

Citation Information

Patent Citations

  • Keyword-based document research hotspot recommending method

    CN106682172A

  • Method and device for determining literature similarity based on semantic analysis

    CN114580557A

  • Universal literature metadata topic analysis method, system and equipment

    CN117493517A

  • Platform content intelligent recommendation method and system based on natural language processing

    CN118551031A

  • Long text information extraction and association analysis method and system based on large model

    CN119761382A

Cited By

  • Composite knowledge chain construction method for scientific research logic representation

    CN120930751A

  • Document deep traceability system based on multi-agent cooperation

    CN121071126A

  • Intelligent research and correlation analysis method and system for scientific research literature based on knowledge graph

    CN122757509A