Short text information processing method and system combined with multiple embedding representations

CN120031041APending Publication Date: 2025-05-23NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510119962.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-23

Smart Images

  • Figure CN120031041A_ABST
    Figure CN120031041A_ABST
Patent Text Reader

Abstract

The invention provides a short text information processing method and system combined with multiple embedding representations, and relates to the technical field of short text processing, and the method comprises the following steps: generating entity embedding vector representations; performing context feature extraction on a target short text through a BERT network, inputting the target short text into a BiGRU network for feature optimization, generating a maximum pooling vector, and generating entity semantic representation; calculating to obtain a target prediction probability; selecting a positive sample and a negative sample to carry out BERT-BiGRU model training; and extracting a CLS vector, a starting position feature vector and an ending position feature vector, performing vector splicing, inputting the vectors into a full-connection neural network layer for vector classification processing, and obtaining a target probability score of the target entity through an activation function. The method solves the technical problem that the traditional entity recognition method mostly depends on vocabulary-level processing and lacks comprehensive understanding of diversified representation of entities, so that the information processing precision is insufficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of short text processing, and in particular to a short text intelligence processing method and system combining multiple embedded representations. Background Art

[0002] In the field of natural language processing, short text intelligence processing and entity recognition have always been important research topics. Short text refers to text with relatively less information and limited contextual information, such as social media posts, messages, reports, news summaries, etc. This type of text is characterized by condensed information and concise expression, but there are also some problems. On the one hand, short texts often cannot provide enough contextual information due to their small amount of information. This makes it difficult for traditional natural language processing methods, especially entity recognition methods based on sequence models, to effectively capture entity relationships and semantic features in texts. Compared with long texts, short texts usually lack sufficient context to help the model understand the precise relationship between words, resulting in poor performance of the model. On the other hand, in short texts, the same entity may have multiple different forms of expression, such as aliases, abbreviations, synonyms, etc., which brings challenges to entity recognition, especially in some specific fields, where entity names often vary. How to accurately identify entities and their specific references has become a difficult point. Summary of the invention

[0003] This application provides a short text intelligence processing method and system combining multiple embedded representations, aiming to solve the technical problem that most traditional entity recognition methods rely on vocabulary-level processing and lack a comprehensive understanding of the diverse representations of entities, resulting in insufficient accuracy in information processing.

[0004] The first aspect disclosed in the present application provides a short text intelligence processing method combined with multiple embedding representations, the method comprising: using multiple embedding representations to process entity name information in an equipment knowledge graph and alias information in an external encyclopedia to generate an entity embedding vector representation; extracting context features from a target short text through a BERT network, and inputting the extracted results into a BiGRU network for feature optimization to generate a maximum pooling vector, and generating an entity semantic representation of a target entity based on the maximum pooling vector, wherein the target entity is a mentioned entity in the target short text; combining the entity semantic representation with the entity embedding vector representation, and performing convolution on the entity semantic representation. The target prediction probability of the target entity is calculated by using a layer, a fully connected layer and a sigmoid activation function; based on the target entity, positive samples are selected from the labeled entities, and negative samples are selected from the candidate entities corresponding to the alias information, and the BERT-BiGRU model is trained in combination with the target prediction probability; the target short text and the target entity are input into the BERT-BiGRU model, the CLS vector, the start position feature vector, and the end position feature vector are extracted, and after vector splicing, they are input into the fully connected neural network layer for vector classification processing, and the target probability score of the target entity is obtained through the activation function.

[0005] The second aspect disclosed in the present application provides a short text intelligence processing system combined with multiple embedding representations, the system is used for the above-mentioned short text intelligence processing method combined with multiple embedding representations, the system includes: an entity embedding vector representation generation module, which is used to use multiple embedding representations to process entity name information in the equipment knowledge graph and alias information in the external encyclopedia to generate entity embedding vector representations; an entity semantic representation generation module, which is used to extract context features from the target short text through the BERT network, and input the extracted results into the BiGRU network for feature optimization to generate a maximum pooling vector, and generate an entity semantic representation of the target entity based on the maximum pooling vector, wherein the target entity is a mentioned entity in the target short text; a target prediction probability calculation module, which is used to convert the entity The semantic representation is combined with the entity embedding vector representation, and is processed through a convolutional layer, a fully connected layer and a sigmoid activation function to calculate the target prediction probability of the target entity; a model training module is used to select positive samples from the labeled entities and negative samples from the candidate entities corresponding to the alias information based on the target entity, and perform BERT-BiGRU model training in combination with the target prediction probability; a target probability score acquisition module is used to input the target short text and the target entity into the BERT-BiGRU model, extract the CLS vector, the start position feature vector, and the end position feature vector, perform vector splicing, and then input them into the fully connected neural network layer for vector classification processing, and obtain the target probability score of the target entity through the activation function.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0007] Through the multiple embedding representation method, it is possible to process information from different data sources, including alias information from equipment knowledge graphs and external encyclopedias, and convert it into entity embedding vector representation. The entity name information in the equipment knowledge graph can provide the model with formal naming and structured semantics about the entity, and the alias information in the external encyclopedia can supplement the different representations of the entity. In this way, the model can not only understand the formal entity name, but also recognize its diverse expressions in different contexts, which enables the system to more comprehensively understand and match the target entity when faced with diverse inputs; through the BERT network for context feature extraction, the model can capture fine-grained context information from the target short text. As a bidirectional pre-trained language model based on Transformer, BERT can integrate the previous and next contexts of each word in the text to provide a more accurate context representation. Then, feature optimization is performed through the BiGRU network to further enhance the sequence modeling capability, so that the model can capture long-term dependencies from the text. Through the maximum pooling operation, the most significant features are retained, thereby ensuring that the model can accurately extract key information when processing short texts; by combining the entity semantic representation with the entity embedding vector representation, a more Comprehensive entity representation, so as to identify and match entities in different contexts. The use of convolutional layers and fully connected layers further enhances the feature extraction capability, so that the model can better classify and predict the target entity. Based on the target entity, positive samples are selected from the labeled entities, and negative samples are selected from the candidate entities corresponding to the alias information. The BERT-BiGRU model is trained in combination with the target prediction probability. The selection of positive and negative samples guides the model to learn how to distinguish the target entity from other irrelevant entities, which enhances the classification accuracy of the model. By guiding the prediction probability, the model can better select samples, reduce errors, and improve the accuracy of entity linking. By inputting the target short text and target entity into the BERT-BiGRU model, extracting the CLS vector, the start position feature vector, and the end position feature vector, and concatenating the vectors, the fully connected neural network is input for classification, and the prediction probability score of the target entity is obtained through the activation function. This process further improves the accuracy of target entity prediction by combining the multi-dimensional features of the target entity. In particular, by concatenating vectors, the model can integrate information from many aspects to achieve more accurate classification and prediction, thereby greatly improving the performance of short text intelligence processing.

[0008] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A flowchart of a short text intelligence processing method combining multiple embedded representations provided in an embodiment of the present application.

[0010] Figure 2 A schematic diagram of the structure of a short text intelligence processing system combined with multiple embedded representations provided in an embodiment of the present application.

[0011] Explanation of the accompanying drawings: entity embedding vector representation generation module 10, entity semantic representation generation module 20, target prediction probability calculation module 30, model training module 40, target probability score acquisition module 50. DETAILED DESCRIPTION

[0012] The embodiments of the present application provide a short text intelligence processing method and system combining multiple embedded representations, thereby solving the technical problem that most traditional entity recognition methods rely on vocabulary-level processing and lack a comprehensive understanding of the diverse representations of entities, resulting in insufficient accuracy in information processing.

[0013] After introducing the basic principles of the present application, various non-limiting implementation methods of the present application will be specifically introduced below in conjunction with the accompanying drawings of the specification. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0014] Embodiment 1, as Figure 1 As shown, the embodiment of the present application provides a short text intelligence processing method combined with multiple embedded representations, the method comprising:

[0015] A multiple embedding representation method is used to process the entity name information in the equipment knowledge graph and the alias information in the external encyclopedia to generate entity embedding vector representation.

[0016] The equipment knowledge graph contains a large number of entities, such as equipment, departments and units, as well as the relationships between them. These entities are represented by triples (entity 1, relationship, entity 2). Entity name information is extracted from the equipment knowledge graph. These entity names are usually the formal names of equipment or items, describing the definition and characteristics of the entity in the knowledge graph. In order to convert these entity name information into vector representation, embedding technology is required. Specifically, the entity names in the knowledge graph will be converted into entity embedding vectors through multiple embedding methods. These vectors are high-dimensional and can capture the semantic and relational characteristics of entities in the graph.

[0017] In addition to the entity names in the equipment knowledge graph, entities may have multiple aliases or synonyms in actual applications. These aliases can usually be found in external encyclopedias. For example, a certain equipment may have different names in different reports, or there may be language differences. The entity alias information in the external encyclopedia is collected. These aliases are usually different expressions of entities, which are used to supplement the information that may be missing in the equipment knowledge graph. Similarly, the alias information is converted into vectors through multiple embedding representations. These vectors will be combined with the entity embedding vectors in the equipment knowledge graph to further enrich the semantic representation of the entity.

[0018] The core of the multiple embedding representation method is to fuse information from different sources (including equipment knowledge graphs and external encyclopedias) to generate a more comprehensive and accurate entity embedding representation. These embedding vectors not only contain the basic characteristics of the entity, but also provide more contextual background in combination with external information, thereby improving the understanding and processing capabilities of the entity.

[0019] The target short text is subjected to context feature extraction through the BERT network, and the extracted result is input into the BiGRU network for feature optimization to generate a maximum pooling vector, and an entity semantic representation of the target entity is generated based on the maximum pooling vector, wherein the target entity is a mentioned entity in the target short text.

[0020] The target short text refers to the part of the text that contains the target entity, such as a piece of text information. In the target short text, the target entity refers to the specific entity mentioned in the text.

[0021] The BERT network is a Transformer-based pre-trained language model that can capture contextual information. The advantage of BERT is that it uses a bidirectional encoding method, which can not only see the information of the previous text, but also the information of the following text, which enables it to fully understand the semantics of each word in the text. The BERT network takes the target short text as input and extracts contextual features in the text through its self-attention mechanism. BERT maps each word or subword in the text to a vector representation and captures the contextual relationship and semantic dependency between words through its multi-layer encoder.

[0022] The BiGRU network is a bidirectional recurrent neural network based on GRU. It combines the GRU gating mechanism and can effectively capture the sequence information in the text. Because it is bidirectional, it can consider the information of the previous and following texts at the same time. This is especially important when processing short texts because short texts may lack sufficient context. The features extracted by the BERT network are input into the BiGRU network for further processing. BiGRU optimizes these features through forward and reverse GRU layers to capture the contextual dependencies of each word or entity in the sequence. The output of the BiGRU layer can better retain the sequence information in the text and dynamically control the information flow through its gating mechanism, so that the model can more accurately understand the semantics of the target entity in the short text.

[0023] In the output processed by the BiGRU network, the maximum pooling operation is used to extract the most important features. The purpose of maximum pooling is to select the maximum value from the output of BiGRU to retain the most significant contextual features in the text. Especially when processing short texts, maximum pooling helps to extract the most representative features.

[0024] The entity semantic representation and the entity embedding vector representation are combined, and processed through a convolutional layer, a fully connected layer and a sigmoid activation function to calculate the target prediction probability of the target entity.

[0025] The entity semantic representation and entity embedding vector are combined by concatenating or fusing their vectors. In this way, a more comprehensive entity representation is obtained, which combines the features extracted from the context and the entity information obtained from the external knowledge base.

[0026] The combined entity representation is input into the convolutional layer. Convolutional neural networks (CNNs) are usually used to process images and sequence data. Here, they are used to extract local features from the combined entity representation. The convolutional layer performs convolution operations on the input features in a sliding window manner to extract discriminative local patterns, which is very useful for improving the accuracy of entity recognition, especially in complex short text intelligence processing.

[0027] The features extracted by the convolutional layer are passed to the fully connected layer. The fully connected layer is responsible for further processing the local features extracted by the convolutional layer, fusing various features, and generating new feature representations through weighted summation of neurons. The fully connected layer can learn higher-level feature representations and can handle more complex patterns.

[0028] After the fully connected layer, an activation function is used for nonlinear transformation. The sigmoid activation function is used here, which compresses the output value between 0 and 1, which makes it suitable for binary classification problems. In this step, the role of the sigmoid activation function is to convert the output of the network into the "existence" probability of the target entity, that is, the predicted probability of whether the target entity is relevant to the given context.

[0029] Through the processing of convolutional layers, fully connected layers and sigmoid activation functions, the network finally outputs the target prediction probability, which represents the relevance of the target entity mentioned in a given target short text in the current context. The closer the probability is to 1, the more relevant the model believes the entity is to the context in the short text; conversely, the closer the probability is to 0, the lower the match between the entity and the short text context.

[0030] Based on the target entity, positive samples are selected from the labeled entities, negative samples are selected from the candidate entities corresponding to the alias information, and BERT-BiGRU model training is performed in combination with the target prediction probability.

[0031] Labeled entities refer to instances that have been clearly labeled as target entities during the training process. Usually, these entities are obtained through manual annotation or through early training data. They have clear category labels or entity information in the dataset. Positive samples refer to entities related to the target entity, such as the alias, specifications, performance, etc. of the target entity. These positive samples have actual semantic or contextual associations with the target entity. The task of the model is to learn how to match the target entity with these positive samples.

[0032] Alias ​​information of entities is usually obtained from external encyclopedias or other knowledge bases. Entities may have multiple aliases, which may refer to the same entity in different contexts. These alias information can help the model identify different representations of the same entity. Candidate entities are candidate lists of entities mentioned by the model in a given target short text. These candidate entities may come from different sources, including entities in knowledge graphs, aliases in encyclopedias, etc. During the training process, candidate entities include not only positive samples related to the target entity, but also entities that may be unrelated to the target entity, i.e., negative samples. Negative samples refer to candidate entities that are unrelated to the target entity. The role of these negative samples is to help the model distinguish entities that are unrelated to the target entity, thereby improving the model's ability to distinguish.

[0033] Combined with the target prediction probability, the model can evaluate the relevance of each candidate entity. Specifically, it uses positive samples selected from labeled entities and negative samples selected from candidate entities to perform training through target prediction probability. The prediction probability of positive samples is usually higher, while the prediction probability of negative samples is lower. The training goal is to optimize the loss function so that the model can correctly distinguish between positive and negative samples, that is, to increase the prediction probability of positive samples and reduce the prediction probability of negative samples.

[0034] The BERT-BiGRU model is a combination of BERT and BiGRU, which enables the model to not only utilize the contextual information extracted by BERT, but also model the sequence information through the BiGRU layer, so as to better understand the entity context in the target short text. The training process adjusts the parameters in the network through the back-propagation algorithm to minimize the difference between the prediction and the actual annotation.

[0035] The target short text and the target entity are input into the BERT-BiGRU model, and the CLS vector, the start position feature vector, and the end position feature vector are extracted. After vector splicing, they are input into the fully connected neural network layer for vector classification processing, and the target probability score of the target entity is obtained through the activation function.

[0036] The target short text and target entity are input into the BERT-BiGRU model, and the CLS (Classification) vector, the start position feature vector, and the end position feature vector are extracted. In the BERT model, the CLS vector is obtained through the [CLS] tag in the model, which is usually used to represent the context information of the entire input text. The CLS vector contains the overall semantic information about the target short text and can represent the global features of the text; the start position feature vector is a feature extracted by the BERT model from the target short text, indicating the start position of the target entity in the text. It helps the model identify the specific position of the target entity in the text and can improve the accuracy of entity positioning; the end position feature vector is similar to the start position feature vector. The end position feature vector indicates the end position of the target entity in the text. Through these two position vectors, the model can accurately capture the boundary of the target entity in the text.

[0037] The extracted CLS vector, start position feature vector and end position feature vector are spliced. The spliced ​​vector will contain the global information of the target short text (from the CLS vector) and the specific position information of the target entity in the text (from the start and end position feature vectors). Through splicing, a comprehensive feature vector containing multiple dimensional information is obtained. This vector reflects the semantic information of the target short text and contains the specific position information of the target entity. The spliced ​​vector is used for subsequent classification processing.

[0038] The concatenated vector will be input into the fully connected neural network layer. The fully connected layer is composed of multiple neurons, each of which is connected to all elements of the input vector. This layer is responsible for further mapping the concatenated feature vector to the relevance prediction space of the target entity. The task of the fully connected layer is to output a prediction result based on the input feature vector, specifically the probability distribution of whether the target entity is in the target short text. The fully connected layer gradually optimizes the prediction ability of the model by learning weights and biases.

[0039] The activation function is used to introduce nonlinear characteristics so that the network can learn complex patterns. Specifically, the sigmoid activation function is used, which maps the output of the fully connected layer to a probability value between 0 and 1. Through the sigmoid function, the model can calculate the probability of whether the target entity matches the target short text. Through the processing of the activation function, a target probability score is generated, which indicates the degree of match between the target entity and the target short text. The higher the score, the stronger the correlation between the target entity and the short text; the lower the score, the lower the match.

[0040] Furthermore, the method of using a multiple embedding representation method to process entity name information in the equipment knowledge graph and alias information in an external encyclopedia to generate an entity embedding vector representation also includes:

[0041] The entity triplets in the equipment knowledge graph are spliced ​​to obtain a graph description text; based on the alias information in the external encyclopedia, the encyclopedia description text is extracted; based on the preset tag constraints of the BERT network, the graph description text and the encyclopedia description text are truncated into long texts according to a preset ratio.

[0042] Entity triples in the equipment knowledge graph are usually expressed as (entity A, relationship, entity B), which describes the relationship between one entity and another, such as entity A belongs to entity B. By concatenating multiple triples, the resulting graph description text will describe the relationship between entities and their attributes in detail. These texts, as descriptions of entities in the knowledge graph, help provide information for subsequent entity embedding and semantic representation.

[0043] An entity may have multiple aliases in different contexts or languages. Based on the alias information in the external encyclopedia, descriptive texts related to the target entity are extracted. These descriptions include the entity's different names, definitions, historical backgrounds, characteristics, etc. By extracting these descriptive texts, we can have a more comprehensive understanding of the entity's multiple definitions and background information.

[0044] The BERT model has a maximum input length limit when processing long texts, usually 512 tokens. If the text exceeds this length, it must be truncated or otherwise processed. The token constraint preset by BERT means that the length of the text input to BERT cannot exceed this maximum limit.

[0045] For description texts extracted from equipment knowledge graphs and external encyclopedias, if their length exceeds the maximum input limit of BERT (512 tokens), the text needs to be truncated according to a preset ratio. This can be achieved by segmenting the text or selecting the first part of the text. Truncation needs to ensure that the most useful part of the text is retained in order to retain important information to the greatest extent. When truncating, the preset ratio is used to determine which parts of the text are retained and which parts are truncated. This ratio is adjusted based on the length of the text, the density of key information, or other strategies to ensure that the most important content is retained.

[0046] Furthermore, the method of generating an entity semantic representation of a target entity based on the maximum pooling vector includes:

[0047] From the output of the BiGRU network, the end position vector of the forward GRU and the start position vector of the reverse GRU are extracted, and vector connection is performed to generate a connection vector; the maximum pooling vector and the connection vector are used as entity semantic representations of the target entity.

[0048] The BiGRU network consists of two independent GRU layers: one is the forward GRU layer and the other is the reverse GRU layer. The forward GRU layer processes the sequence in chronological order, while the reverse GRU layer processes the sequence in reverse chronological order. The output of the forward GRU network generates a vector, which represents the state of the current time step after processing each time step of the input sequence. The end position vector of the forward GRU refers to the output vector of the last time step in the sequence, which represents all the contextual information from the beginning of the text to a certain position of the target entity; the output of the reverse GRU network is similar, but its processing order is from the end of the text forward. The start position vector of the reverse GRU refers to the output vector of the first time step in the sequence, which contains the contextual information from the target entity to the end of the text.

[0049] The end position vector extracted from the forward GRU and the start position vector extracted from the reverse GRU represent the context information at both ends of the target entity in the sequence. These two vectors are connected, usually element-by-element, and the resulting connection vector contains the bidirectional context information of the target entity, which can more comprehensively represent the semantics of the entity.

[0050] Furthermore, the BERT network is based on a bidirectional Transformers structure, expressed as follows:

[0051] Query(Q)=XW Q ;

[0052] Key(K)=XW K ;

[0053] Value(V)=XW V ;

[0054]

[0055] Where X is the input vector, W Q , W K , W V It is the weight matrix learned by the model, which is used to convert the input vector into Query, Key, and Value;

[0056] Among them, Q, K, V represent query, key and value matrices, QK T Represents the dot product of the transposed Query matrix and Key matrix, d K Represents the dimension of the Key vector, and softmax() is used to constrain the sum of weights to 1.

[0057] Specifically, the basic structure of the BERT network is derived from the encoder part of the Transformers model, which uses the self-attention mechanism to capture the contextual dependencies in the text, which is mathematically expressed as follows:

[0058] Query(Q)=XW Q ;

[0059] Key(K)=XW K ;

[0060] Value(V)=XW V ;

[0061]

[0062] For a given input sequence x, x represents the text data input into the BERT network. Its three transformations Query, Key, and Value are initially calculated through linear transformations. In the above equations, Q, K, and V are query, key, and value matrices derived from the input data. Finally, the attention score is calculated and normalized. Overall, the query, key, and value matrices interact with each other through the self-attention mechanism to determine which parts of the information are most relevant in the context. The correlation between the query and the key is calculated by dot product, and the result is applied to the value matrix. Finally, the contextual representation of each word in the sentence is calculated. This mechanism enables BERT to take into account all contextual information when understanding the text, not just the order from left to right or right to left, which is crucial for capturing complex semantic relationships.

[0063] Furthermore, the BiGRU network includes an update gate and a reset gate for controlling the information flow inside the GRU unit.

[0064] The BiGRU network is a variant of the Recurrent Neural Network (RNN), which is used to solve the problem of gradient disappearance or increase in extended sequences. BiGRU combines the features of the Gated Recurrent Unit (GRU) and the bidirectional RNN, and has great advantages in processing sequence text processing tasks. The BiGRU model contains two types of gate mechanisms: update gates and reset gates, which control the information flow inside the GRU unit.

[0065] Furthermore, the mathematical formula of a GRU unit is as follows:

[0066] z t =σ(W z ·[h t-1 ,x t ]);

[0067] r t =σ(W r ·[h t-1 ,x t ]);

[0068] Among them, z t is the update gate at time step t, r t is the reset gate at time step t, σ is the sigmoid activation function, and W z , W r They are the matrix weights, h t-1 is the hidden state of the previous time step, x t is the input for the current time step.

[0069] Specifically, a GRU unit can be expressed mathematically as follows:

[0070] z t =σ(W z ·[h t-1 ,x t ]);

[0071] r t =σ(W r ·[h t-1 ,x t ]);

[0072] The update gate determines how much of the information in the current time step comes from the information in the previous time step. The update gate controls how the hidden state is updated. t is close to 1, which means that most of the information comes from the hidden state of the previous time step; if z t The closer it is to 0, the more information comes from the current input.

[0073] The reset gate determines how much the information of the current time step is independent of the hidden state of the previous time step. It controls to what extent the state of the previous time step should be reset. t Close to 0, it means that the previous hidden state has less influence and the model will rely more on the current input; if r t If it is close to 1, the hidden state of the previous time step has a greater impact on the current calculation.

[0074] Furthermore, the BiGRU network consists of two independent GRU layers, wherein the forward GRU processes the sequence in chronological order and the reverse GRU processes the sequence in reverse chronological order.

[0075] BiGRU combines the features of gated recurrent units and bidirectional RNNs. "Bidirectional" means that the model can process text information in both forward and backward directions. Specifically, BiGRU consists of two independent GRU layers: one processes the sequence in chronological order (forward GRU), and the other processes the sequence in reverse chronological order (reverse GRU). The output of each time step is the connection of the outputs from the forward and backward GRUs, thereby incorporating past and future contexts in each time step.

[0076] In summary, the short text intelligence processing method combined with multiple embedded representations provided in the embodiments of the present application has the following technical effects:

[0077] Through the multiple embedding representation method, it is possible to process information from different data sources, including alias information from equipment knowledge graphs and external encyclopedias, and convert it into entity embedding vector representation. The entity name information in the equipment knowledge graph can provide the model with formal naming and structured semantics about the entity, and the alias information in the external encyclopedia can supplement the different representations of the entity. In this way, the model can not only understand the formal entity name, but also recognize its diverse expressions in different contexts, which enables the system to more comprehensively understand and match the target entity when faced with diverse inputs; through the BERT network for context feature extraction, the model can capture fine-grained context information from the target short text. As a bidirectional pre-trained language model based on Transformer, BERT can integrate the previous and next contexts of each word in the text to provide a more accurate context representation. Then, feature optimization is performed through the BiGRU network to further enhance the sequence modeling capability, so that the model can capture long-term dependencies from the text. Through the maximum pooling operation, the most significant features are retained, thereby ensuring that the model can accurately extract key information when processing short texts; by combining the entity semantic representation with the entity embedding vector representation, a more Comprehensive entity representation, so as to identify and match entities in different contexts. The use of convolutional layers and fully connected layers further enhances the feature extraction capability, so that the model can better classify and predict the target entity. Based on the target entity, positive samples are selected from the labeled entities, and negative samples are selected from the candidate entities corresponding to the alias information. The BERT-BiGRU model is trained in combination with the target prediction probability. The selection of positive and negative samples guides the model to learn how to distinguish the target entity from other irrelevant entities, which enhances the classification accuracy of the model. By guiding the prediction probability, the model can better select samples, reduce errors, and improve the accuracy of entity linking. By inputting the target short text and target entity into the BERT-BiGRU model, extracting the CLS vector, the start position feature vector, and the end position feature vector, and concatenating the vectors, the fully connected neural network is input for classification, and the prediction probability score of the target entity is obtained through the activation function. This process further improves the accuracy of target entity prediction by combining the multi-dimensional features of the target entity. In particular, by concatenating vectors, the model can integrate information from many aspects to achieve more accurate classification and prediction, thereby greatly improving the performance of short text intelligence processing.

[0078] Embodiment 2 is based on the same inventive concept as the short text intelligence processing method combined with multiple embedded representations in the previous embodiment. Figure 2 As shown, the embodiment of the present application provides a short text intelligence processing system combined with multiple embedded representations, the system comprising:

[0079] The entity embedding vector representation generation module 10 is used to process the entity name information in the equipment knowledge graph and the alias information in the external encyclopedia using a multiple embedding representation method to generate an entity embedding vector representation; the entity semantic representation generation module 20 is used to extract context features from the target short text through the BERT network, and input the extracted results into the BiGRU network for feature optimization to generate a maximum pooling vector, and generate an entity semantic representation of the target entity based on the maximum pooling vector, wherein the target entity is a mentioned entity in the target short text; the target prediction probability calculation module 30 is used to combine the entity semantic representation with the entity embedding vector representation, and through the convolution layer, the fully connected layer and the sig The target prediction probability of the target entity is calculated by the moid activation function; the model training module 40 is used to select positive samples from the labeled entities and negative samples from the candidate entities corresponding to the alias information based on the target entity, and perform BERT-BiGRU model training in combination with the target prediction probability; the target probability score acquisition module 50 is used to input the target short text and the target entity into the BERT-BiGRU model, extract the CLS vector, the start position feature vector, and the end position feature vector, perform vector splicing, and input them into the fully connected neural network layer for vector classification processing, and obtain the target probability score of the target entity through the activation function.

[0080] Furthermore, the system further includes a long text truncation module to perform the following operation steps:

[0081] The entity triplets in the equipment knowledge graph are spliced ​​to obtain a graph description text; based on the alias information in the external encyclopedia, the encyclopedia description text is extracted; based on the preset tag constraints of the BERT network, the graph description text and the encyclopedia description text are truncated into long texts according to a preset ratio.

[0082] Furthermore, the system further includes an entity semantic representation acquisition module to perform the following operation steps:

[0083] From the output of the BiGRU network, the end position vector of the forward GRU and the start position vector of the reverse GRU are extracted, and vector connection is performed to generate a connection vector; the maximum pooling vector and the connection vector are used as entity semantic representations of the target entity.

[0084] Furthermore, the BERT network is based on a bidirectional Transformers structure, expressed as follows:

[0085] Query(Q)=XW Q ;

[0086] Key(K)=XW K;

[0087] Value(V)=XW V ;

[0088]

[0089] Where X is the input vector, W Q , W K , W V It is the weight matrix learned by the model, which is used to convert the input vector into Query, Key, and Value. Q, K, and V represent the query, key, and value matrices. QK T Represents the dot product of the transposed Query matrix and Key matrix, d K Represents the dimension of the Key vector, and softmax() is used to constrain the sum of weights to 1.

[0090] Furthermore, the BiGRU network includes an update gate and a reset gate for controlling the information flow inside the GRU unit.

[0091] Furthermore, the mathematical formula of a GRU unit is as follows:

[0092] z t =σ(W z ·[h t-1 ,x t ]);

[0093] r t =σ(W r ·[h t-1 ,x t ]);

[0094] Among them, z t is the update gate at time step t, r t is the reset gate at time step t, σ is the sigmoid activation function, and W z , W r They are the matrix weights, h t-1 is the hidden state of the previous time step, x t is the input for the current time step.

[0095] Furthermore, the BiGRU network consists of two independent GRU layers, wherein the forward GRU processes the sequence in chronological order and the reverse GRU processes the sequence in reverse chronological order.

[0096] Through the above detailed description of the short text intelligence processing method combined with multiple embedded representations in this specification, those skilled in the art can clearly understand the short text intelligence processing system combined with multiple embedded representations in this embodiment. Since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.

[0097] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A short text intelligence processing method combining multiple embedding representations, characterized in that: The method comprises: The multiple embedding representation method is used to process the entity name information in the equipment knowledge graph and the alias information in the external encyclopedia to generate the entity embedding vector representation; Extract context features from the target short text through the BERT network, input the extracted results into the BiGRU network for feature optimization, generate a maximum pooling vector, and generate an entity semantic representation of the target entity based on the maximum pooling vector, wherein the target entity is a mentioned entity in the target short text; The entity semantic representation and the entity embedding vector representation are combined, and processed through a convolution layer, a fully connected layer and a sigmoid activation function to calculate the target prediction probability of the target entity; Based on the target entity, select positive samples from the labeled entities, select negative samples from the candidate entities corresponding to the alias information, and perform BERT-BiGRU model training in combination with the target prediction probability; The target short text and the target entity are input into the BERT-BiGRU model, and the CLS vector, the start position feature vector, and the end position feature vector are extracted. After vector splicing, they are input into the fully connected neural network layer for vector classification processing, and the target probability score of the target entity is obtained through the activation function.

2. The short text intelligence processing method combining multiple embedded representations as claimed in claim 1, characterized in that: The method further includes: processing entity name information in the equipment knowledge graph and alias information in the external encyclopedia using a multiple embedding representation method to generate an entity embedding vector representation. splicing entity triplets in the equipment knowledge graph to obtain graph description text; Extracting encyclopedia description text based on the alias information in the external encyclopedia; Based on the preset tag constraints of the BERT network, the long text of the graph description text and the encyclopedia description text is truncated according to a preset ratio.

3. The short text intelligence processing method combining multiple embedded representations as claimed in claim 1, characterized in that: The method of generating an entity semantic representation of a target entity based on the maximum pooling vector comprises: Extracting the end position vector of the forward GRU and the start position vector of the reverse GRU from the output of the BiGRU network, and performing vector connection to generate a connection vector; The maximum pooling vector and the connection vector are used as entity semantic representations of the target entity.

4. The short text intelligence processing method combining multiple embedded representations as claimed in claim 1, characterized in that: The BERT network is based on a bidirectional Transformers structure, expressed as follows: Query(Q)=XW Q ; Key(K)=XW K ; Value(V)=XW V ; Where X is the input vector, W Q , W K , W V It is the weight matrix learned by the model, which is used to convert the input vector into Query, Key, and Value; Among them, Q, K, V represent query, key and value matrices, QK T Represents the dot product of the transposed Query matrix and Key matrix, d K Represents the dimension of the Key vector, and softmax() is used to constrain the sum of weights to 1.

5. The short text intelligence processing method combining multiple embedded representations as claimed in claim 1, characterized in that: The BiGRU network includes an update gate and a reset gate for controlling the information flow inside the GRU unit.

6. The short text intelligence processing method combining multiple embedded representations as claimed in claim 5, characterized in that: The mathematical formula of a GRU unit is as follows: Among them, z t is the update gate at time step t, r t is the reset gate at time step t, σ is the sigmoid activation function, and W z , W r They are the matrix weights, h t-1 is the hidden state of the previous time step, x t is the input for the current time step.

7. The short text intelligence processing method combining multiple embedded representations as claimed in claim 1, characterized in that: The BiGRU network consists of two independent GRU layers, wherein the forward GRU processes the sequence in chronological order and the reverse GRU processes the sequence in reverse chronological order.

8. A short text intelligence processing system combining multiple embedded representations, characterized in that: A system for implementing the short text intelligence processing method combined with multiple embedded representations as described in any one of claims 1 to 7, comprising: An entity embedding vector representation generation module is used to process entity name information in the equipment knowledge graph and alias information in the external encyclopedia using a multiple embedding representation method to generate entity embedding vector representation; An entity semantic representation generation module is used to extract context features from the target short text through a BERT network, and input the extracted results into a BiGRU network for feature optimization to generate a maximum pooling vector, and generate an entity semantic representation of the target entity based on the maximum pooling vector, wherein the target entity is a mentioned entity in the target short text; A target prediction probability calculation module is used to combine the entity semantic representation and the entity embedding vector representation, and process them through a convolution layer, a fully connected layer and a sigmoid activation function to calculate the target prediction probability of the target entity; A model training module, for selecting positive samples from the labeled entities based on the target entity, selecting negative samples from the candidate entities corresponding to the alias information, and performing BERT-BiGRU model training in combination with the target prediction probability; The target probability score acquisition module is used to input the target short text and the target entity into the BERT-BiGRU model, extract the CLS vector, the start position feature vector, and the end position feature vector, perform vector splicing, and input them into the fully connected neural network layer for vector classification processing, and obtain the target probability score of the target entity through the activation function.