AI medical comment hierarchical classification method and system based on hyperbolic space attention

By introducing the hyperbolic spatial attention mechanism into AI medical reviews and combining it with BERT and LSTM, the shortcomings of traditional methods in hierarchical modeling are addressed, achieving higher accuracy and robustness in classification, which is suitable for text classification in various medical scenarios.

CN120744115APending Publication Date: 2025-10-03SHANDONG NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510931688.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing AI hierarchical classification methods for medical reviews lack the ability to model hierarchical structures. Traditional attention mechanisms have difficulty capturing the implicit semantic hierarchy and structural nesting in review texts in Euclidean space, resulting in poor classification accuracy and result interpretation capabilities.

Method used

The hyperbolic space attention mechanism is introduced to integrate the semantic modeling capabilities of BERT, the sequence modeling advantages of LSTM, and the hierarchical expression capabilities of hyperbolic space. The semantic enhancement features are calculated through the attention mechanism in hyperbolic geometric space to achieve accurate hierarchical classification of AI medical reviews.

Benefits of technology

It improves the classification accuracy and robustness of AI medical reviews, can better understand the semantic information of complex hierarchical structures, is suitable for text classification tasks in various medical scenarios, and has strong portability and versatility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744115A_ABST
    Figure CN120744115A_ABST
Patent Text Reader

Abstract

The invention discloses an AI medical comment hierarchical classification method and system based on hyperbolic space attention, and relates to the technical field of natural language processing, and the method comprises the steps: obtaining a to-be-classified original AI medical comment text, inputting the preprocessed AI medical comment text into a BERT model, and carrying out the coding and modeling of a multi-layer Transform self-attention mechanism, thereby obtaining an AI medical comment hierarchical classification model. Extracting context-sensitive semantic features; inputting the semantic features into a bidirectional LSTM network, capturing a long dependency relationship and grammatical logic in the semantic features, and extracting semantic enhancement features; based on a hyperbolic space attention module, generating attention weights according to structural distances among the semantic enhancement features in hyperbolic geometry, and calculating to obtain the semantic enhancement features subjected to attention weight weighted aggregation; and performing classification based on the weighted and aggregated features, and outputting a multi-level category label of the text. According to the method, accurate hierarchical classification and semantic aggregation of the AI medical comments can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to an AI medical review hierarchical classification method and system based on hyperbolic spatial attention. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the rapid development of information technology, artificial intelligence (AI) has become a core force for change across all industries, especially in the medical field, where its application potential is enormous. With the aging of the global population, the increasing number of patients with chronic diseases, and the inequitable distribution of medical resources, the application of AI in the medical field is seen as a vital tool for addressing these challenges. For example, AI can help build intelligent health monitoring systems, using wearable devices to collect patient physiological data such as heart rate, blood pressure, and sleep quality. Deep learning algorithms can be used to analyze this data in real time, promptly identifying potential health risks and issuing warnings, enabling early intervention for diseases. Alternatively, AI can develop personalized treatment and rehabilitation plans for patients based on large amounts of patient data.

[0004] Although some progress has been made in the research and development of AI medical technology, there is still a lack of systematic and in-depth analysis of the development status of AI medical care and its influencing factors. The hierarchical classification method for AI medical review data faces challenges, including: (1) Lack of hierarchical structure modeling capabilities: AI medical review texts often have obvious hierarchical semantic structures. For example, a review may simultaneously involve three different levels of content: "diagnosis and treatment process", "doctor's attitude" and "treatment effect". These contents are not parallel, but have implicit subordination and abstraction levels. However, most existing methods only consider flat labels when classifying, and fail to effectively utilize the implicit hierarchical semantics or structural differences in the reviews. This results in limited recognition capabilities for multi-dimensional hierarchical labels such as technology maturity and ethical acceptance, and an inability to effectively extract the implicit hierarchical structure information in the text, resulting in poor classification accuracy and result interpretation capabilities.

[0005] (2) The model’s expressive power is limited by Euclidean space: Traditional attention mechanisms or classification embeddings have good performance in capturing semantic weights, but they are all constructed based on Euclidean space, that is, the construction process relies on Euclidean distance or dot product operations, making it difficult to effectively capture the implicit semantic hierarchy, structural nesting, or nonlinear gradients in the review text, and difficult to finely model hierarchical semantics in geometric space, which affects the final classification accuracy. Summary of the Invention

[0006] To address the deficiencies of the above-mentioned prior art, the present invention provides a hierarchical classification method and system for AI medical reviews based on hyperbolic space attention. By introducing hyperbolic geometry space, integrating the hyperbolic representation with the attention mechanism, and constructing a hyperbolic attention mechanism suitable for the medical context, the present invention forms a classification method that integrates the semantic modeling capabilities of BERT, the sequence modeling advantages of LSTM, and the hierarchical expression capabilities of hyperbolic space, thereby achieving accurate hierarchical classification and semantic aggregation of AI medical reviews, and providing key technical support for AI-assisted medical services.

[0007] In a first aspect, the present invention provides an AI medical review hierarchical classification method based on hyperbolic spatial attention.

[0008] An AI-based hierarchical classification method for medical reviews based on hyperbolic spatial attention, including: Obtain the original AI medical review text to be classified and perform standardized preprocessing on the obtained text; The pre-processed AI medical review text is input into the BERT model, which is then encoded and modeled using a multi-layer Transformer self-attention mechanism to extract context-sensitive semantic features. Input semantic features into a bidirectional LSTM network to capture long-term dependencies and grammatical logic in semantic features and extract semantic enhancement features; Based on the hyperbolic spatial attention module, attention weights are generated according to the structural distance between semantic enhancement features in hyperbolic geometry, and semantic enhancement features after weighted aggregation by attention weights are calculated; Classification is performed based on the semantically enhanced features after weighted aggregation, and multi-level category labels of the original AI medical review text are output.

[0009] In the second aspect, the present invention provides an AI medical review hierarchical classification system based on hyperbolic spatial attention.

[0010] An AI medical review hierarchical classification system based on hyperbolic spatial attention, including: The data acquisition and preprocessing module is used to obtain the original AI medical review text to be classified and perform standardized preprocessing on the obtained text; The encoding module is used to input the pre-processed AI medical review text into the BERT model, and extract context-sensitive semantic features through encoding and modeling using a multi-layer Transformer self-attention mechanism; The semantic sequence modeling module is used to input semantic features into the bidirectional LSTM network, capture the long-term dependencies and grammatical logic in the semantic features, and extract semantic enhancement features; A hyperbolic attention module is used to generate attention weights based on the structural distance between semantic enhancement features in hyperbolic geometry based on the hyperbolic spatial attention module, and calculate the semantic enhancement features after weighted aggregation of the attention weights; The classification prediction module is used to perform classification based on the semantically enhanced features after weighted aggregation and output multi-level category labels for the original AI medical review text.

[0011] In a third aspect, the present invention also provides an electronic device comprising: a memory for storing executable instructions; and a processor for implementing the above-mentioned AI medical review hierarchical classification method based on hyperbolic spatial attention when executing the executable instructions stored in the memory.

[0012] In a fourth aspect, the present invention also provides a computer-readable storage medium storing executable instructions for causing a processor to execute the executable instructions to implement the above-mentioned AI medical review hierarchical classification method based on hyperbolic spatial attention.

[0013] In a fifth aspect, the present invention also provides a computer program product, which includes executable instructions, and the executable instructions are stored in a computer-readable storage medium; wherein, when the processor of the electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the above-mentioned AI medical review hierarchical classification method based on hyperbolic spatial attention is implemented.

[0014] One or more of the above technical solutions have the following beneficial effects: 1. The present invention proposes a hierarchical classification method and system for AI medical reviews based on hyperbolic space attention, introduces hyperbolic geometric space, integrates the hyperbolic representation with the attention mechanism, and constructs a hyperbolic attention mechanism suitable for the medical context. On this basis, BERT, LSTM and the hyperbolic attention mechanism are organically integrated, fully combining BERT's powerful modeling capabilities in semantic understanding, LSTM's sensitivity to sequence-dependent information and the advantages of hyperbolic space in processing hierarchical structures, to achieve a deeper and more structured understanding of AI medical reviews, and to achieve accurate hierarchical classification and semantic aggregation of AI medical reviews, providing key technical support for AI-assisted medical services. Compared with traditional classification methods based on flat feature spaces or relying only on a single semantic model, the present invention shows higher robustness and accuracy in processing AI medical reviews with ambiguous semantics, overlapping concepts and complex label hierarchies.

[0015] 2. The hyperbolic attention mechanism constructed in this invention can automatically focus on the core concepts and their semantic levels in medical texts, adjust the attention weight by spatial structural distance, effectively reduce redundant information interference, and improve the generalization ability and structural adaptability of the classification model.

[0016] 3. The method proposed in the present invention can construct a semantic hierarchy without relying on an external knowledge graph, has strong portability and versatility, is suitable for text classification tasks in a variety of medical scenarios, and has significant practical value and promotion prospects.

[0017] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0019] Figure 1 Flowchart of the AI ​​medical review hierarchical classification method based on hyperbolic spatial attention in an embodiment of the present invention; Figure 2 This is an analysis flow chart of hierarchical classification based on hyperbolic space and LSTM in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] It should be noted that the following detailed descriptions are exemplary only and are intended to describe specific embodiments and provide further explanation of the present invention, and are not intended to limit the exemplary embodiments according to the present invention. Unless otherwise indicated, all technical and scientific terms used herein have the same meanings as those commonly understood by those of ordinary skill in the art to which the present invention belongs. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0021] Technical term explanation: BERT (Bidirectional Encoder Representations from Transformers): a pre-trained language representation model.

[0022] LSTM (Long Short-Term Memory): An improved recurrent neural network with a three-part structure consisting of a forget gate, an input gate, and an output gate. It can effectively capture long-distance dependency information and avoid the vanishing gradient problem.

[0023] Example 1 In response to the problems of the lack of hierarchical modeling capabilities and the limitation of model expression capabilities to Euclidean space in the existing hierarchical classification methods of AI medical reviews, and considering that hyperbolic geometry has been proven to be more suitable for representing hierarchical data in the field of natural language processing, such as syntax trees and knowledge graphs, hyperbolic space provides exponential expansion capabilities, which can more effectively embed hierarchical information and improve the distinction between semantic expression and clustering. To this end, this embodiment introduces hyperbolic geometry into the hierarchical classification task of AI medical review texts, integrates the hyperbolic representation with the attention mechanism, and constructs a hyperbolic attention mechanism suitable for the context of AI medical reviews. On this basis, a new classification method is proposed that can integrate the BERT semantic modeling capabilities, the LSTM sequence modeling advantages and the hierarchical expression capabilities of hyperbolic space, to achieve accurate hierarchical classification and semantic aggregation of AI medical reviews, and provide key technical support for the research of AI medical services.

[0024] The AI ​​medical review hierarchical classification method based on hyperbolic spatial attention proposed in this embodiment is as follows: Figure 1 As shown, specifically including: Obtain the original AI medical review text to be classified and preprocess the obtained text; The pre-processed AI medical review text is input into the BERT model, which is then encoded and modeled using a multi-layer Transformer self-attention mechanism to extract context-sensitive semantic features. Input semantic features into a bidirectional LSTM network to capture long-term dependencies and grammatical logic in semantic features and extract semantic enhancement features; Based on the hyperbolic spatial attention module, attention weights are generated according to the structural distance between semantic enhancement features in hyperbolic geometry, and semantic enhancement features after weighted aggregation by attention weights are calculated; Classification is performed based on the semantically enhanced features after weighted aggregation, and multi-level category labels of the original AI medical review text are output.

[0025] The following content provides a detailed introduction to the AI ​​medical review hierarchical classification method proposed in this embodiment.

[0026] First, before performing AI medical review hierarchical classification, a multi-level medical semantic classification model is constructed based on the BERT model, bidirectional LSTM network, hyperbolic spatial attention module and classification module. The trained model is used to perform the above hierarchical classification. The training process of the multi-level medical semantic classification model is as follows: Figure 2 As shown, Step S1: Construct an AI medical review text dataset. Specifically, the process involves designing and implementing a large-scale unlabeled text corpus collection and standardized preprocessing process for medical artificial intelligence applications. Several unlabeled AI medical review texts to be analyzed are obtained from medical review websites and SCI-indexed websites to form an unlabeled text dataset. The obtained texts are then preprocessed and hierarchical modeling of AI medical review data features is performed to construct the AI ​​medical review text dataset.

[0027] Step S2: Design an attention mechanism based on hyperbolic geometric space (i.e., hyperbolic attention mechanism), and construct a hyperbolic attention module based on the hyperbolic attention mechanism to enhance the model's ability to model hierarchical semantic structures in medical texts.

[0028] Step S3: Construct a multi-level medical semantic classification model that integrates BERT semantic representation, LSTM temporal dependency modeling, and hyperbolic attention mechanism.

[0029] Step S4: Use the AI ​​medical review text dataset to perform end-to-end training on the model based on a unified training framework, and use a multi-index system to evaluate accuracy, robustness, and generalization capabilities.

[0030] In this embodiment, unlabeled text is first collected from a medical review platform and an SCI literature database, and high-quality corpus preprocessing is achieved through cleaning, denoising, and BERT-compatible word segmentation. Then, an attention mechanism based on hyperbolic space is constructed, and the hyperbolic distance is replaced by Euclidean similarity to enhance the model's ability to capture complex hierarchical semantics. BERT, LSTM, and hyperbolic attention are then integrated to design a multi-level, multi-label classification model to improve semantic understanding and classification accuracy. Finally, an end-to-end training strategy is adopted, combined with an optimizer and learning rate scheduling, to conduct a comprehensive multi-indicator evaluation to ensure model stability and generalization ability to meet the task requirements of AI medical review.

[0031] In step S1, the corresponding data set is obtained according to the analysis purpose, and articles, comments and SCI journals from multiple online forums are crawled to obtain several AI medical review texts, which are then aggregated into an AI medical review text data set.

[0032] Taking the development of AI medical care as an example, a Python-based data crawler program can be used to collect articles and comments about the application of AI in the medical field from multiple medical website forums and included SCI papers, and then organize and summarize the data into an AI medical review dataset.

[0033] Based on the systematic collection and analysis of relevant reviews, literature, and social feedback texts on the application of artificial intelligence in medical scenarios, a hierarchical, multi-dimensional, and clearly structured modeling framework is established to quantitatively characterize the development stage, social acceptance, and institutional environment adaptation of AI medical technology in real-world environments. This modeling process not only serves the assessment of technology maturity, but also further introduces two external dimensions: ethics and law, thereby realizing a comprehensive semantic labeling system with a "technology-ethics-law" three-way coupling, effectively improving the abstract modeling capabilities of complex review corpora. Specifically, it includes: Step S11: Use Python to write a web crawler program to automatically collect raw data from multiple sources, including articles, reviews, and research abstracts related to the application of artificial intelligence in the medical field from mainstream medical review websites (such as Dingxiangyuan, Haodaifu Online, and Zhihu Medical Topic Zone) and SCI literature databases (such as PubMed and Web of Science). These raw texts constitute a preliminary data set, denoted as , where each Represents a piece of raw text data.

[0034] In order to facilitate the subsequent topic modeling and sentiment analysis, the original text data is first structured and the text is divided into title, author or publisher identification, analysis object text content, keywords, release time and collection timestamp, that is, each data It can be represented as the following structured tuple: (1) in, Indicates the title, Identify the author or publisher, The text content of the main analysis object is It is the keyword that comes with the extraction or crawling. is the time the article or comment was published, It is the timestamp collected by the system.

[0035] Step S12: Keyword filtering. Specifically, based on the constructed keyword set related to AI medical care , filter out or A subset of the data , that is, only retain text data related to "AI+medical".

[0036] Step S13: Text cleaning. Specifically, for The text content in , perform operations such as unified character encoding, remove HTML tags, emoticons, meaningless words (such as stop words), and perform word segmentation. The processed dataset is recorded as: (2) in, Represents a cleaned and normalized text structure.

[0037] Step S14, semantic standardization. Considering the diversity of semantic expressions in Chinese, there are cases where the keywords are the same but the semantic directions are different, such as "AI doctor mistakes" and "AI doctor efficiency is high" both contain the keyword "AI doctor". To this end, similarity analysis based on contextual semantic vectors is introduced, such as using the BERT model to generate sentence vectors. , in order to judge the meaning of the sentence, screen the semantically consistent texts, and manually remove or semantically cluster the samples with large semantic deviations to enhance the consistency of the data.

[0038] Through the above time division, the final unlabeled text corpus is , which can be regarded as a high-quality AI medical text review dataset, which will provide a solid foundation for subsequent topic modeling, sentiment analysis, influencing factor identification, etc. The dataset is in the form of: (3) in, For the text content after standardization, This is structured metadata information, including time, keywords, etc. The entire process complies with the preprocessing process in unsupervised learning scenarios, ensuring the integrity, relevance, and quality of the corpus, laying a data foundation for research on the development path of AI medical care.

[0039] Step S15, hierarchical modeling of AI medical review data. Specifically, for the unlabeled AI medical review text corpus obtained after preprocessing, the modeling adopts a three-dimensional main variable system: technical maturity, ethical acceptance, and legal and policy maturity, and sets three levels (Level 1-3) for each dimension from the three development stages of the initial conception period, the system verification period, and the formal application period, so as to depict the challenges and social reactions faced by AI medical care at different stages. Based on the different levels of different dimensions, the AI ​​medical review text is hierarchically labeled and modeled. This hierarchical design facilitates the model to further perform multivariate regression, path analysis or sentiment tendency association research. In addition, Level 0 is set to indicate that a single data lacks information on a certain dimension, which is specifically described as: (1) Technology maturity Level 0 information is missing.

[0040] Level 1: The initial stage. AI is only in the theoretical exploration and concept building stage, with no concrete implementation. Research tends to focus on expert opinions and descriptions of potential value.

[0041] Level 2 Functional Verification Phase: Algorithm solutions, technical papers, and prototype systems have been produced. Key modules are feasible, but full deployment has not yet begun.

[0042] Level 3: System Integration and Implementation. The AI ​​system has completed clinical testing or deployment and has entered the formal use process in medical institutions, demonstrating a high level of technical maturity.

[0043] (2) Ethical acceptance Level 0 information is missing.

[0044] Level 1: Initial awareness stage. Society and researchers have yet to systematically understand ethical issues. Ethical discussions remain at the conceptual level, with no technical implementation measures.

[0045] Level 2: Mechanism Embedding Stage. Ethical factors such as data privacy, informed consent, and explainability modules have been considered within the system, and a preliminary ethics testing mechanism has been introduced.

[0046] Level 3: Social Acceptance and Compliance. The system has passed ethical compliance certification (such as HIPAA and GDPR), is widely accepted by patients and medical staff, and has mature ethical mechanisms.

[0047] (3) Legal and policy maturity Level 0 information is missing.

[0048] Level 1: Unregulated. AI healthcare has not yet been incorporated into the legal framework. Discussions are only at the strategic level, and legal responsibilities and regulatory scope are unclear.

[0049] Level 2: Pilot and Policy Phase. Regulatory sandboxes and draft regulations are introduced, and the government or institutions issue industry guidance policies. Some pilot projects begin to operate in compliance.

[0050] Level 3: Regulations implementation and enforcement. AI medical products have been formally incorporated into the national regulatory system, requiring approval and licensing, with a complete compliance path and accountability mechanism.

[0051] In step S2, the attention mechanism is a key technology for simulating "selective attention" in human cognition. It assigns weights by calculating the correlation between inputs, thereby enhancing the model's ability to express important information. Traditional attention mechanisms generally rely on distance or dot product calculations in Euclidean space, which are often inadequate when processing data with hierarchical structures, exponential growth, or nonlinear distributions. To this end, this embodiment proposes a hyperbolic attention mechanism (HyperbolicAttention), which performs attention calculations in hyperbolic space. The key to constructing the hyperbolic attention mechanism lies in its calculation and representation update of attention weights in hyperbolic space. Compared with traditional Euclidean space, hyperbolic space has the geometric property of negative curvature, which makes it more advantageous in representing tree-like, nested, or multi-level hierarchical data. In natural language processing, many language structures such as lexical dependencies, syntactic levels, and semantic categories naturally have hierarchical nested relationships. Therefore, hyperbolic space provides a more appropriate geometric representation for modeling such complex structures. Its core construction is to replace the Euclidean metric with hyperbolic distance to construct a more expressive attention weight function under non-Euclidean structure.

[0052] Furthermore, from the perspective of theoretical basis and geometric construction, the hyperbolic attention mechanism is introduced, which specifically includes the following steps: Step S21: Mapping the input feature vector to the unit hyperbolic space based on the mapping function; Step S22: Calculate the hyperbolic distance between the eigenvectors based on a hyperbolic distance function that conforms to geometric axioms; Step S23: constructing Softmax attention weight based on hyperbolic distance; Step S24: perform weighted aggregation on the input feature vector using the attention weight to obtain the final attention-weighted feature vector.

[0053] In the above step S21, it is assumed that the input is a set of vectors, each of which is initially in Euclidean space. Since the hyperbolic attention mechanism requires distance calculation in hyperbolic space, these Euclidean vectors need to be mapped to the unit hyperbolic space, which represents the hyperbolic space as a unit open sphere. In order to complete the mapping, the exponential mapping function in hyperbolic geometry can be used, which is defined as follows: (4) In the above formula, represents the input feature vector, The above mapping ensures that the input vector falls within the unit sphere and retains the directionality and partial modulus information. It should be noted that for simplicity, the embedded vector is still recorded as , but assumes it is in hyperbolic space.

[0054] In step S22, a hyperbolic distance function is constructed, which is also the basis of hyperbolic attention. Based on this hyperbolic distance function, the hyperbolic distance between feature vectors can be calculated. Traditional attention mechanisms rely on dot products or Euclidean distances, but the distance between vectors in hyperbolic space has a different definition. In the spherical model, the hyperbolic distance between two points is defined as: (5) The above distance definition reflects the geometric properties of hyperbolic space: as a point approaches the boundary on the sphere, the distance grows much faster than in Euclidean space. This property makes hyperbolic space very suitable for modeling tree structures, hierarchical relationships, or data structures with exponential growth, especially in the fields of medicine, graph structures, and text hierarchical modeling. The numerator of this distance function is the square of the Euclidean difference, and the denominator is the "residual distance" from each point to the boundary of the unit sphere. The entire ratio reflects the geometric tension between two points in a space with negative curvature. The function is used to map the result to a real distance scale.

[0055] In step S23, by The idea is to set the weight of the point with smaller distance to be larger, that is, the attention is more focused on the vector closer to the current point. The attention weight is defined as follows: (6) in, Is the hyperbolic distance metric, representing two vectors The hyperbolic distance between , which measures the difference between two vectors in hyperbolic geometric space. This weight expression has the same normalized structure as the traditional attention mechanism, but due to the different distance metric, it produces a nonlinear enhancement effect. In hyperbolic space, the exponential term of distant points decays more rapidly due to the rapid growth of distance. This "selective attention" effectively avoids interference from noise or irrelevant nodes while maintaining computational stability, thereby improving the model's representation capabilities for sparse or complex data structures.

[0056] In step S24, similar to traditional attention, after completing the attention weight construction, all input vectors are weighted summed to construct the output vector. In this embodiment, an approximate linear combination method is used to complete the aggregation, which can be expressed as: (7) In the above formula, It represents the hyperbolic attention output built around the i-th vector, and realizes the vector expression of complex structural relationships by integrating the attention weights of other points.

[0057] In this mechanism, the spatial structure is first defined by a selected hyperbolic geometric model (such as the Poincare disk model or the Poincare half-plane model), and the relationship between any two points is measured by the hyperbolic distance, which reflects the relative position of words or sentences at the semantic level. Figure 1 As shown in the figure, the function F is used to map the input vector from Euclidean space to hyperbolic space, while its derivative F' is responsible for the numerical conversion between spaces when calculating gradients or performing backpropagation. This model can more reasonably distribute and calculate attention weights while efficiently maintaining the language hierarchy, thereby improving the attention mechanism's ability to model complex language phenomena. Ultimately, this structure significantly enhances the model's expressiveness and performance in processing natural language tasks with hierarchical features.

[0058] In step S3, the pre-processed AI medical review text is input into a multi-level medical semantic classification model for classification. The processing flow is the process of the AI ​​medical review hierarchical classification method based on hyperbolic spatial attention proposed in this embodiment, which is as follows: First, the input raw text undergoes standardization preprocessing, including removing special symbols, unifying capitalization, and cleaning noise. This ensures the accuracy of subsequent word segmentation and vectorization, as well as the standardization and consistency of data in the subsequent modeling phase. Subsequently, the entire model adopts the BERT+LSTM+hyperbolic attention mechanism fusion architecture. BERT's word segmenter is used to segment the text into subword units that match its vocabulary. Special tags such as [CLS] are added to represent the overall sentence meaning and [SEP] is used to identify sentence boundaries. Three types of embeddings are generated simultaneously: word embedding (Token Embedding), position embedding (Position Embedding), and segment embedding (Segment Embedding), which together constitute the standard input tensor. BERT is used as a pre-trained bidirectional language model based on the Transformer architecture. The multi-head self-attention mechanism is used to model the bidirectional dependency between any two words in the input sequence, obtaining context-sensitive high-dimensional semantic vectors. These contextual embeddings are encoded and modeled by the multi-layer Transformer self-attention mechanism and serve as the deep foundational features of semantic expression, namely, context-sensitive semantic features. To further enhance the model's ability to model language structure and word order dependencies, BERT's output is fed into an LSTM (Long Short-Term Memory) network. LSTM is a recurrent neural network with forget, input, and output gates. It excels at capturing long-term dependencies and grammatical logic information in time series. Its introduction complements BERT's weakness in temporal modeling, enabling the model to possess both global semantic parsing capabilities and the ability to understand local language fluidity. The contextually enhanced features output by the LSTM are then fed into a hyperbolic spatial attention module, which generates attention weights based on the structural distances between the semantically enhanced features in hyperbolic geometry. This module then calculates the semantically enhanced features after weighted aggregation using the attention weights. This module migrates the traditional attention mechanism from Euclidean space to hyperbolic geometric space to better characterize the implicit hierarchical structure, nested relationships, and semantic gradients in medical text. In hyperbolic space, the exponential growth of inter-node distances makes it more suitable for representing the nonlinear relationships between label hierarchies and complex medical concepts, thereby improving the model's structural sensitivity and expression accuracy in multi-level labeling scenarios (e.g., levels 0 to 3).

[0059] Finally, classification is performed based on the semantically enhanced features after weighted aggregation, and multi-level category labels of the original AI medical review text are output.

[0060] In addition, taking into account the problem that the number of level 0 samples in the training set is much higher than that of other levels, this embodiment specifically designs an adaptive corrective loss function (AdaptiveCorrective Loss) for dealing with situations where level 0 features have a high proportion, in order to improve the robustness and generalization ability of the model under label imbalance. The above model further introduces an adaptive corrective loss function, which dynamically adjusts the gradient weights of samples of different levels, strengthens the recognition tendency of medium and high-level samples, and suppresses the overfitting behavior of the model on the main frequency labels in the early stage of training. Through this loss design, the model effectively reduces the level confusion rate while maintaining the overall accuracy, and enhances the multi-level classification ability in complex AI medical review texts. In summary, the model designed in this embodiment realizes the triple information fusion of semantics, timing and structure in structure, enhances the adaptability to unbalanced labels in optimization strategy, and has strong semantic expression and practical transferability as a whole.

[0061] Specifically, in step S31, the input AI medical review text is first preprocessed, including word segmentation, adding special tags such as CLS and SEP, and mapping it to the corresponding word ID sequence through the vocabulary. After BERT tokenization, the input sequence is in the form of: (8) In the above formula, Indicates the i The vocabulary index of the word.

[0062] Step S32: Input the preprocessed AI medical review text into the BERT model based on the multi-layer Transformer encoder. Through the multi-layer Transformer self-attention mechanism, the input sequence is encoded and modeled to extract context-sensitive semantic features.

[0063] The above step S32 specifically includes: Step S321: Input to the BERT model based on the multi-layer Transformer encoder, and record the sequence as , No. i Word embedding vector representation of position go through L Layer encoding, output context-aware vector, is: (9) in, d represents the word vector dimension, L Represents a Transformer layer.

[0064] Step S322: In each layer of Transformer, the calculation of the self-attention mechanism is: given the input matrix , that is, The output of the layer, calculate the query, key, and value matrix, and we can get: (10) in, is the learned weight matrix, The attention dimension.

[0065] Step S323: Calculate the attention score as follows: (11) The final output is: (12) Step S324: After multi-layer encoding, the vector sequence H output by BERT integrates global context information. BERT uses a multi-layer Transformer architecture to output the context representation as: (13) Among them, each , For the i The contextual embedding vector of each word, d Output dimension for BERT.

[0066] Step S33: Enter LSTM, the recursive relationship is as follows: (14) (15) (16) (17) (18) (19) in, is the current time step input, is the hidden state of the previous step, is the memory unit state, 、 is the weight matrix of each gate, is the bias term of each gate, σ is activation function, is the hyperbolic tangent function, ⊙ is (Element-wise) product.

[0067] The final output is , represents the temporal dependency encoding vector, the As the input of the next step hyperbolic attention mechanism.

[0068] Step S34: performing a hyperbolic attention calculation operation in the hyperbolic space, including: Step 341: Output of the previous step One-to-one mapping is the input of this step , as shown in Equation 6, the relationship between each pair of vectors is calculated by hyperbolic distance.

[0069] Step S342: Calculate the hyperbolic distance as shown in Formula 5. .

[0070] Step S343: As shown in Formula 7, the weights calculated using hyperbolic attention in non-Euclidean space are used to weight the input vector .

[0071] Step S35: During the training process, the imbalanced modeling level in the dataset is taken into consideration, and an adaptive correction loss function based on dynamic weight reconstruction and inter-category comparison regularization is adopted. At the same time, a multi-stage parameter optimization strategy is used for iterative training to continuously iteratively optimize the model parameters.

[0072] Specifically, due to the high sparsity and unbalanced feature distribution of the data used in this embodiment, especially the presence of a large number of data samples containing only partial features in multiple feature classification dimensions, conventional classification models face significant challenges during the training process. Among them, those missing items that do not show specific category information under a certain feature classification are uniformly marked as "level 0" to indicate that they are in a category-less state. However, this label definition causes serious category imbalance problems in the actual training process: on the one hand, the proportion of level 0 samples in the overall training set is extremely high, which easily leads to its dominant gradient contribution in the loss function, thereby causing the model to overfit on category-less data; on the other hand, the number of level 1, 2, and 3 samples that truly have clear classification significance is relatively small in the data, and it is difficult for the model to learn enough discriminative features from them during the optimization process, thereby inhibiting the effective modeling ability of key categories and reducing the overall classification accuracy and generalization ability.

[0073] To address the training bottleneck caused by such extreme label imbalance, this embodiment designs an adaptive correction loss function for the case of high-proportion level 0 features. This loss function integrates two key mechanisms: dynamic weight rebalancing and inter-class contrast regularization. The former adjusts the weight contribution of level 0 samples in the total loss in real time based on the prior probability of sample distribution and the uncertainty of model prediction, thereby preventing them from having a dominant effect in gradient propagation; the latter enhances the model's discriminative learning ability for minority classes by introducing positive and negative sample comparison terms between levels 1, 2, and 3, while maintaining the stability of level 0 classification and improving the recognition sensitivity and boundary clarity of valid categories.

[0074] Specifically, in step S351, dynamic weights are introduced to level 0 and normal levels 1, 2, and 3 respectively to reduce the contribution of level 0 to the total loss, which can be expressed as: (20) Among them, N is the number of samples, which represents the total number of samples in the data set; i is the index of the sample, ranging from 1 to N; k is the index of the category, representing each category in the classification problem, For category k The weight can be used to adjust the importance of each category. It is usually used when the categories are unbalanced. Fewer categories can be given higher weights to balance the loss function. is the indicator function, indicating that if the sample i The true label Equal to category k,but , otherwise it is 0. By designing this function, it can be used to select and accumulate samples that match the current category; Representation category k The logarithm of the predicted probability is used to calculate the cross entropy loss.

[0075] Step S352: Design the weights of level 0 and normal level. The weight of level 0 is lower to avoid its dominant loss, which can be expressed as: (twenty one) in, It represents the dynamic adjustment weight coefficient of level 0 samples, and P is the proportion probability of samples of a certain level in the entire data set, such as Indicates the probability of the proportion of level 0 samples in the entire data set; using To alleviate the impact of excessive P(y=0), the weight is ultimately kept between 0 and 1, and the larger P(y=0), The smaller it is, the more effectively it can reduce the weight contribution of level 0 samples in the loss function.

[0076] The weights for normal levels are: (twenty two) in, Indicates the k The weight coefficient of the class-level label is inversely proportional to the weight of the normal level and the sample distribution, which enhances the learning of minority classes; is a small positive number used to smooth the denominator; Indicates the probability of the proportion of samples of level k in the entire data set.

[0077] Step S353: Introduce inter-class comparison regularization and construct a regularized loss function to avoid the model's excessive focus on level 0 while improving the ability to distinguish levels 1, 2, and 3. It can be expressed as: (twenty three) in, Is an adjustable regularization strength hyperparameter, which determines the weight of this regularization term in the total loss function. The larger the value, the stronger the suppression effect. It represents the average denominator after summing the regularization terms, which prevents the regularization term value from expanding or shrinking due to the difference in the number of samples; and These two parameters together constitute the scaling factor of the regularization loss, Indicates the regularization strength of the level 0 sample set, which is used to control the suppression of level 0.

[0078] Step S354: Combining the above two parts, the final loss function is obtained: (twenty four) In step S4, to ensure that the hierarchical classification model based on BERT, LSTM, and hyperbolic spatial attention mechanism proposed in this embodiment has good generalization ability and semantic expression performance in practical applications, the training and evaluation phase is systematically divided into three key sub-steps: Step S41, training mechanism and parameter optimization strategy; Step S42: Model evaluation method and performance measurement indicators; Step S43: ablation experiment and generalization verification.

[0079] In step S41, to ensure the stability and efficiency of the model in semantic representation and hierarchical label discrimination, the training mechanism adopts a multi-stage parameter optimization strategy. Combining the characteristics of the pre-trained language model and the embedding properties of the hyperbolic geometric structure, the following training process is designed: In step S411, the initialization strategy uses pre-trained BERT parameters as the initialization weights for the semantic encoder to maintain consistency in semantic understanding and domain generalization. The remaining modules (such as the LSTM and HyperbolicAttention layers) use a hybrid of He initialization and normal distribution initialization to enhance gradient propagation stability.

[0080] Step S412: For optimizer selection, AdamW optimizer is selected and the weight decay coefficient is set to 0.01 to avoid overfitting, especially in the scenario of label imbalance in processing medical corpus.

[0081] In step S413, the learning rate scheduling mechanism adopts the warm up-linear decay strategy, which first linearly increases the learning rate by 10% warm up step, and then linearly decays it until the end of training. This strategy can effectively improve the convergence speed and stability of the model in the early stage of training.

[0082] Step S414: A gradient clipping mechanism is introduced to avoid gradient explosion. The clipping threshold is set to 1.0, and a gradient monitoring mechanism is set for the attention layer and LSTM layer in the model.

[0083] Step S415, mixed precision training (FP16), is used to accelerate training and save video memory, and is particularly suitable for large-scale corpus training scenarios.

[0084] In step S42, in order to comprehensively evaluate the classification accuracy and hierarchical perception ability of the model in the medical review corpus, this embodiment designs a multi-dimensional, structured performance evaluation framework that combines traditional classification indicators and hierarchical indicators. The indicators include Accuracy, Precision, Recall and Macro-F1 to measure the overall classification performance, especially the average performance in the case of multiple labels and uneven data distribution.

[0085] In step S43, in order to systematically evaluate the impact and contribution of each module of the proposed model structure on the overall performance, multiple sets of structural ablation studies are constructed in this stage. The purpose is to explore the marginal gain of each structural unit on the model's predictive ability and expressive ability through module stripping and replacement operations. The designed control experiments include the following four types of modified models: (1) Hyperbolic Attention replacement experiment: the hyperbolic space attention mechanism (HyperbolicAttention) in the original model is replaced with traditional Euclidean Attention to observe the impact of embedding geometric space on the ability to capture hierarchical dependencies; (2) Sequence modeling elimination experiment: the LSTM module is removed, and only the BERT semantic representation result is retained and directly fed into the classifier for decision making, aiming to analyze the role of time series structure modeling (Sequential DependencyModeling) in contextual reasoning; all experiments use the same training hyperparameter configuration (including Warmup-Linear Decay learning rate scheduling, AdamW optimizer, Early Stopping mechanism, etc.).

[0086] Experimental results show that the complete model has a 4.7% improvement in Macro-F1 on high-level labeled samples (especially level 2 and level 3), and the level-wise misclassification rate is reduced by an average of 0.8, showing stronger fine-grained hierarchical recognition capabilities compared to degenerate structures; the hyperbolic attention mechanism significantly improves the model's expressiveness and attention focus stability in processing entity nested structures, long-distance dependencies, and semantically dense paragraphs (such as complex term definitions and disease descriptions); maintaining good F1 performance under category imbalance indicates that the model has strong hierarchical concept fitting capabilities and structural generalization capabilities in non-Euclidean space. In addition, the complete model shows consistent cross-context transfer robustness on multiple datasets, indicating that it has good adaptability and generalizability in real-world heterogeneous medical corpus. The above ablation results fully demonstrate the key role and system synergy of the core modules in the model proposed in this embodiment, especially the hyperbolic space structure modeling, in the high-dimensional AI medical review semantic task.

[0087] Example 2 This embodiment provides an AI medical review hierarchical classification system based on hyperbolic spatial attention, including: The data acquisition and preprocessing module is used to obtain the original AI medical review text to be classified and perform standardized preprocessing on the obtained text; The encoding module is used to input the pre-processed AI medical review text into the BERT model, and extract context-sensitive semantic features through encoding and modeling using a multi-layer Transformer self-attention mechanism; The semantic sequence modeling module is used to input semantic features into the bidirectional LSTM network, capture the long-term dependencies and grammatical logic in the semantic features, and extract semantic enhancement features; A hyperbolic attention module is used to generate attention weights based on the structural distance between semantic enhancement features in hyperbolic geometry based on the hyperbolic spatial attention module, and calculate the semantic enhancement features after weighted aggregation of the attention weights; The classification prediction module is used to perform classification based on the semantically enhanced features after weighted aggregation and output multi-level category labels for the original AI medical review text.

[0088] Example 3 This embodiment provides an electronic device, including: a memory for storing executable instructions; and a processor for implementing the above method provided in this embodiment when executing the executable instructions stored in the memory.

[0089] Example 4 This embodiment further provides a computer-readable storage medium storing executable instructions. When the executable instructions are executed by a processor, the processor will be caused to execute the above method provided in this embodiment.

[0090] Example 5 This embodiment provides a computer program product including executable instructions, which are computer instructions stored in a computer-readable storage medium. When a processor of an electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the electronic device performs the method provided in this embodiment.

[0091] The steps involved in the above embodiments 2 to 5 correspond to those in embodiment 1. For detailed implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media that includes one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to perform any method of the present invention.

[0092] Those skilled in the art will appreciate that the modules or steps of the present invention described above can be implemented using a general-purpose computer device. Alternatively, they can be implemented using program code executable by a computing device, which can then be stored in a storage device and executed by the computing device. Alternatively, they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0093] The above description is only a preferred embodiment of the present invention. Although the specific implementation of the present invention is described in conjunction with the accompanying drawings, it does not limit the scope of protection of the present invention. Those skilled in the art should understand that on the basis of the technical solution of the present invention, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present invention.

Claims

1. An AI medical review hierarchical classification method based on hyperbolic spatial attention, characterized by: include: Obtain the original AI medical review text to be classified and perform standardized preprocessing on the obtained text; The pre-processed AI medical review text is input into the BERT model, which is then encoded and modeled using a multi-layer Transformer self-attention mechanism to extract context-sensitive semantic features. Input semantic features into a bidirectional LSTM network to capture long-term dependencies and grammatical logic in semantic features and extract semantic enhancement features; Based on the hyperbolic spatial attention module, attention weights are generated according to the structural distance between semantic enhancement features in hyperbolic geometry, and semantic enhancement features after weighted aggregation by attention weights are calculated; Classification is performed based on the semantically enhanced features after weighted aggregation, and multi-level category labels of the original AI medical review text are output.

2. The AI ​​medical review hierarchical classification method based on hyperbolic spatial attention according to claim 1, characterized in that: Based on the BERT model, bidirectional LSTM network, hyperbolic spatial attention module, and classification module, a multi-level medical semantic classification model is constructed. The training process of the multi-level medical semantic classification model is as follows: Obtain a number of unlabeled AI medical review texts, preprocess and perform hierarchical modeling on the obtained AI medical review texts, and construct an AI medical review text dataset; A multi-level medical semantic classification model was trained using an AI medical review text dataset. Taking into account the imbalanced modeling levels in the dataset, an adaptive modified loss function based on dynamic weight reconstruction and inter-category comparison regularization was adopted. A multi-stage parameter optimization strategy was used for iterative training to continuously optimize model parameters. The model training is completed through continuous iterative training until the loss function is minimized or the set number of iterations is reached.

3. The AI ​​medical review hierarchical classification method based on hyperbolic spatial attention as claimed in claim 2, characterized in that: The pretreatment includes: Each AI medical review text is structured and divided into title, author or publisher identifier, analysis object text content, keywords, publication time, and collection timestamp; Based on a preset set of AI medical-related keywords, the structured AI medical review text is screened; Perform text cleaning on the screened AI medical review texts, perform unified character encoding on the analysis object text content, remove HTML tags, emoticons, and meaningless words, and perform word segmentation; Based on the similarity of contextual semantic vectors, semantically consistent texts are screened, and semantic standardization of all texts is completed to obtain an unlabeled AI medical review text corpus.

4. The AI ​​medical review hierarchical classification method based on hyperbolic spatial attention according to claim 2, characterized in that: The hierarchical modeling of several AI medical review texts is as follows: For the unlabeled AI medical review text corpus obtained after preprocessing, a three-dimensional main variable system of technology maturity, ethical acceptance, and legal and policy maturity is adopted, and three levels are set for each dimension from the three development stages of initial conception, system verification, and formal application. Based on the different levels of different dimensions, the AI ​​medical review text is hierarchically annotated; among them, level 0 is also set to indicate that a single data lacks information on a certain dimension.

5. The AI ​​medical review hierarchical classification method based on hyperbolic spatial attention according to claim 1, characterized in that: The hyperbolic spatial attention mechanism adopted by the hyperbolic spatial attention module is: Based on the mapping function, the input feature vector is mapped to the unit hyperbolic space; Based on the hyperbolic distance function, the hyperbolic distance between feature vectors is calculated; Construct Softmax attention weights based on hyperbolic distance, perform weighted aggregation on the input feature vectors with attention weights, and obtain the final attention-weighted feature vectors; The mapping function is: ; The hyperbolic distance function is: ; In the above formula, represents the input feature vector, Represents the feature vector after mapping.

6. The AI ​​medical review hierarchical classification method based on hyperbolic spatial attention according to claim 2, characterized in that: The construction of an adaptive correction loss function based on dynamic weight reconstruction and inter-category comparison regularization includes: Dynamic weights are introduced for level 0 and normal level respectively, and the weights of level 0 and normal level are designed. The loss function based on dynamic weight reconstruction is constructed as follows: ; In the above formula, , Indicates the dynamic adjustment weight coefficient for level 0 samples, , Indicates the k The weight coefficient of the class level label, P is the proportion probability of samples of a certain level in the entire data set, is a small positive number, is the indicator function, which means that if the sample i The true label Equal to category k ,but , otherwise 0; Representation category k The predicted probability of Introduce inter-category contrast regularization and construct the regularized loss function: ; In the above formula, is the regularization strength of the sample set at level 0, is an adjustable regularization strength hyperparameter, represents the average denominator after summing the regular terms; The loss function based on dynamic weight reconstruction and the regularization loss function are combined to obtain the final adaptive correction loss function.

7. An AI medical review hierarchical classification system based on hyperbolic spatial attention, characterized by: include: The data acquisition and preprocessing module is used to obtain the original AI medical review text to be classified and perform standardized preprocessing on the obtained text; The encoding module is used to input the pre-processed AI medical review text into the BERT model, and extract context-sensitive semantic features through encoding and modeling using a multi-layer Transformer self-attention mechanism; The semantic sequence modeling module is used to input semantic features into the bidirectional LSTM network, capture the long-term dependencies and grammatical logic in the semantic features, and extract semantic enhancement features; A hyperbolic attention module is used to generate attention weights based on the structural distance between semantic enhancement features in hyperbolic geometry based on the hyperbolic spatial attention module, and calculate the semantic enhancement features after weighted aggregation of the attention weights; The classification prediction module is used to perform classification based on the semantically enhanced features after weighted aggregation and output multi-level category labels for the original AI medical review text.

8. An electronic device, characterized in that: include: a memory for storing executable instructions; A processor, configured to implement the AI ​​medical review hierarchical classification method based on hyperbolic spatial attention as described in any one of claims 1 to 6 when executing the executable instructions stored in the memory.

9. A computer-readable storage medium, characterized in that Executable instructions are stored, which are used to cause the processor to execute the executable instructions to implement the AI ​​medical review hierarchical classification method based on hyperbolic spatial attention as described in any one of claims 1-6.

10. A computer program product, characterized in that The computer program product includes executable instructions stored in a computer-readable storage medium; When the processor of the electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the AI ​​medical review hierarchical classification method based on hyperbolic spatial attention described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • CRM sales verbal skill intelligent analysis method based on natural language processing

    CN121390085A

  • Large model data grading method and device

    CN121681711A