Short text semantic enhancement method based on domain knowledge graph and satellite image
By combining domain knowledge graphs and satellite images, the semantic information of short text is enhanced, and the problems of insufficient information and lack of context in the recognition of short text naming entity are solved, achieving higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510156404.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-17
AI Technical Summary
When dealing with short text named entity recognition, the prior art faces the problems of insufficient information and lack of context, resulting in insufficient recognition accuracy and robustness.
Short text semantic enhancement method based on domain knowledge graph and satellite images is used to retrieve triples from domain knowledge graph through large language model auxiliary labeling, multiple recall and rearrangement methods, and provide background knowledge with satellite images, and perform statement processing and entity labeling.
It significantly improves the accuracy and robustness of short text entity recognition, reduces the difficulty of the model when facing short text data, and enhances the robustness and generalization capabilities of the system.
Smart Images

Figure CN120163158A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a short text semantic enhancement method based on a domain knowledge graph and satellite images. Background Art
[0002] In the field of natural language processing, the context association of information plays a crucial role in understanding text semantics. Specifically in the named entity recognition task, by introducing more context information, the model can more accurately identify the key entities and their categories in the text. This is because humans often rely on the surrounding environment and the information before and after to infer the meaning of information when communicating. Similarly, the model also needs to rely on context information to parse the accurate meaning of entities when processing text.
[0003] Especially for shorter texts, they usually contain relatively less information, which makes the model face greater challenges in processing short text entity recognition. For example, in the sentence "The Adora Magic City left Qingdao", the short text only provides a basic event framework - "The Adora Magic City" left "Qingdao". For people who are not familiar with cruise ships, "The Adora Magic City" here may be misinterpreted as the code name of a train or a bus, and "Qingdao" may be simply understood as the name of a city rather than a port name. However, the correct entities are: "The Adora Magic City" here refers to a cruise ship, and "Qingdao" is a famous port in China. Therefore, how to enrich the context information has become the key to improving the accuracy of short text entity recognition.
[0004] To enrich the context information of short texts, the following three methods are usually adopted in the prior art for semantic enhancement.
[0005] (1) Using voice information to enhance text semantics;
[0006] (2) Using text structure information to enhance text semantics;
[0007] (3) Using picture information to enhance text semantics.
[0008] Using voice information to enhance text semantic understanding is a potential method, which improves the accuracy of entity recognition to a certain extent. However, this method still faces technical challenges, and its effects may vary in different scenarios and tasks.
[0009] Using text structure information to enhance text semantic understanding can effectively process and integrate the connection between words to a certain extent, but it still seems powerless in dealing with the context missing problem of short sentences or broken sentences and is difficult to perfectly solve the understanding problem of short sentences in the absence of context.
[0010] Enhancing text semantics using image information can significantly improve the expressive effect of text. The image provides a visual supplement that can assist readers in constructing a richer mental image and help form a more vivid and comprehensive scene picture in the brain. However, the information provided by the image is more of a "vague outline" rather than an accurate description of entities. For information that requires in-depth understanding of specific entity attributes or locations, relying solely on the description of the image may be limited. Summary of the Invention
[0011] To address the problem of short text named entity recognition, the present invention proposes a short text semantic enhancement method that adopts a comprehensive strategy of "short text + domain knowledge graph + satellite image". This method can provide richer and multi-level semantic information for the entity recognition process by integrating the advantages of different information sources, thereby significantly reducing the recognition difficulty of the model when facing short text data and simultaneously significantly improving the robustness and accuracy of the entire system.
[0012] The present application provides a short text semantic enhancement method based on a domain knowledge graph and satellite images, including:
[0013] S1: Use a large language model to assist in annotating entities and their types in the short text;
[0014] S2: Adopt multi-way recall and re-ranking means to retrieve domain knowledge graph triples associated with the short text from the domain knowledge graph, and use the triple after sentence transformation as entity knowledge in short text entity recognition;
[0015] S3: Determine whether the short text of the original corpus contains an attached satellite image: If the short text of the original corpus has an attached satellite image, directly use the attached satellite image; otherwise, search for a satellite image associated with the short text in the satellite image library as background knowledge for short text entity recognition;
[0016] S4: Perform sentence transformation on the triple and the satellite image respectively to unify the input form and obtain the text after semantic enhancement;
[0017] S5: Use entity knowledge and background knowledge to perform entity annotation on the text after semantic enhancement, give the true entity label, and obtain the annotated corpus.
[0018] According to the method provided by the present application, wherein step S2 includes:
[0019] S21: Preprocess the original short text, and the preprocessing includes keyword extraction and text vectorization to obtain keywords and short text vectors;
[0020] S22: Recall a series of candidate subgraphs from the domain knowledge graph in a multi-way manner and retrieve a set of triples associated with the short text;
[0021] S23: Vectorize the triples in the triple set using a short text vectorization model, calculate the cosine distance between each triple and the short text vector as the semantic relevance score, sort them in descending order of the semantic relevance score, select the top N triples, sentence the top N triples, and use the sentenced N triples as entity knowledge in short text entity recognition.
[0022] According to the method provided by this application, wherein step S3 includes:
[0023] S31: Segment and vectorize the short text;
[0024] S32: Conduct multi-way retrieval in the satellite image library to recall several satellite images related to the short text;
[0025] S33: Convert the satellite image into text using a judgment model and perform vectorization processing on the text to obtain a satellite image vector, and then calculate the cosine distance between the satellite image vector and the short text vector as the relevance score between the satellite image and the short text;
[0026] S34: Sort the scores from high to low, and select the satellite image with the highest score as background knowledge for short text entity recognition.
[0027] According to the method provided by this application, for annotating a corpus, the corpus includes:
[0028] Short text;
[0029] Entities and their types included in the short text;
[0030] Domain knowledge graph triples associated with the short text;
[0031] Satellite images associated with the short text.
[0032] According to the method provided by this application, in step S4,
[0033] The sentence processing rule for triples is: The {attribute / relationship name} of {entity name} is {attribute value / relationship value};
[0034] The sentence processing rule for satellite images is: Use models such as image interpretation and object detection to extract information such as location, space, object, layout, and environment from the satellite image, and summarize the information as the background information of the short text.
[0035] This application also provides a short text understanding model, and the model architecture of this model includes:
[0036] A bidirectional pre-trained language model for feature extraction and fusion, trained based on large-scale unannotated text, captures context information and semantic features, and provides semantic representations for subsequent processing;
[0037] A bidirectional long short-term memory network for considering the forward and backward dependencies of sequential data;
[0038] Conditional random field.
[0039] This application also provides a training method for a short text understanding model. Using the labeled corpus obtained by the above semantic enhancement method as the training corpus for model training, the training method includes:
[0040] S1: Split the labeled corpus into a training set, a test set, and a validation set;
[0041] S2: Set the key hyperparameters of the model;
[0042] S3: Start iterative training, and gradually optimize the short text understanding model under the guidance of the hyperparameters to reach the optimal performance state.
[0043] According to the method provided by this application, the ratio of the training set, the test set, and the validation set is set to 3:1:1.
[0044] According to the method provided by this application, the training set is used for model training, the test set is used to verify the generalization ability of the model, and the validation set is used to select the optimal model parameters during the model adjustment stage.
[0045] According to the method provided by this application, the key hyperparameters of the model include the learning rate and the regularization coefficient.
[0046] This technology combines a domain knowledge graph and satellite imagery to reduce the difficulty of the model in understanding short texts. Among them, the domain knowledge graph can provide word-level semantic information. Compared with pictures, satellite imagery contains richer information, such as location information, spatial information, layout information, etc. Therefore, relying on the joint enhancement of the domain knowledge graph and satellite imagery, this technology can not only enrich the background information of short texts but also strengthen the model's understanding of specific entities, not only reducing the training difficulty of the short text entity recognition model but also ensuring the accuracy of the model in the production environment. Brief Description of the Drawings
[0047] The following will further illustrate the above characteristics, technical features, advantages and their implementation manners of this application in a clear and understandable manner through the description of preferred embodiments and in conjunction with the drawings. The following drawings are only intended to make a schematic illustration and explanation of this application, and do not limit the scope of this application. Among them:
[0048] Figure 1 Is the overall process of the technical solution provided by this application;
[0049] Figure 2 Schematic diagram of the triple retrieval process for the domain knowledge graph;
[0050] Figure 3 Schematic diagram of the satellite image retrieval process. Detailed implementation manners
[0051] For a clearer understanding of the technical features, objectives, and effects of the present application, the detailed implementation manners of the present application will now be described with reference to the accompanying drawings.
[0052] In order to enrich the context information of short texts, the following three methods are usually adopted in the prior art for semantic enhancement.
[0053] (1) Using voice information to enhance text semantics
[0054] In the field of natural language processing, the interpretation of text often depends on accurately identifying entities and their boundaries. The Chinese language is characterized by no spaces between words, which poses a challenge to the establishment of entity boundaries. To overcome this difficulty, people have started to explore multi-modal assisted methods. As an effective auxiliary means, the voice modality provides an additional dimension for text analysis by capturing the pronunciation features of language information. When processing text containing entity information, the voice modality can identify the natural pauses between different words during speech, thus helping the model more accurately locate the start and end positions of entities. This cross-modal information fusion not only improves the accuracy of entity recognition but also helps to more accurately parse the text structure in the absence of clear interval information, thereby enhancing the overall performance of natural language processing tasks. Therefore, integrating the voice modality into the text processing flow has become one of the key strategies for effectively improving the accuracy of entity boundary recognition in scenarios where there is a lack of clear text interval information.
[0055] (2) Using text structure information to enhance text semantics
[0056] The glyph structure of Chinese characters plays a crucial role in Chinese text processing and entity recognition. By parsing and identifying the radicals or structural components of Chinese characters, rich context information can be provided for entity recognition, thus enhancing the accuracy and efficiency of recognition. A key method is to utilize the subordination relationship. By analyzing the connection mode of each Chinese character with its basic radical, the overall composition of the Chinese character can be understood, thereby assisting in the localization and recognition of entities. In addition, introducing traditional Chinese characters as the research object can more comprehensively cover the diversity of Chinese texts because traditional Chinese characters often contain more radical and structural information than simplified Chinese characters. To effectively extract structural features from Chinese characters, convolutional neural networks (CNNs) are widely applied in this field. CNNs can efficiently process image data. By performing convolutional operations on the images of Chinese characters, the local structural features inside the Chinese characters and the spatial relationships between different radicals can be captured. Such a processing method can not only accurately identify the radical composition of Chinese characters but also effectively distinguish the subtle differences between different Chinese characters, providing strong feature support for subsequent entity recognition tasks. In practical applications, by combining the subordination relationship of Chinese characters and traditional Chinese character resources, as well as using the method of extracting radical features by CNNs, the accuracy and efficiency of Chinese entity recognition can be significantly improved. This method is not only applicable to the automatic analysis of documents and ancient books but also can play an important role in modern text processing, information retrieval, and natural language processing, further promoting the research and development of artificial intelligence in the field of language understanding and processing.
[0057] (3) Use picture information to enhance text semantics
[0058] When constructing complex natural language processing tasks, social media text information often suffers from quality deficiencies for various reasons. These deficiencies may manifest as grammar errors, inaccuracies in expression, noise information, and a lack of sufficient context information, making it difficult to directly extract effective and in-depth information from the text. However, by combining the use of image information, the ability to understand and predict text semantics can be effectively enhanced. Images, especially those associated with the text content, can provide intuitive, specific, and often entity information that cannot be directly conveyed by words. These images can supplement the missing details in the text, correct grammar errors, reduce information uncertainty, and even provide clear visual evidence when the text description is vague or ambiguous. For example, a picture containing a product model can help consumers understand the actual appearance of the product when the text description is incorrect or incomplete, thus reducing doubts in the purchase decision-making process. In addition, images can capture non-verbal cues and emotional information, which is crucial for understanding the implicit semantics in the text. For instance, emojis, the expressions, postures, and background elements of the people in the picture may contain emotional, attitudinal, or situational information related to the text content, which is difficult to fully express clearly through the text. By combining this image information, the accuracy of the model in tasks such as sentiment analysis, event understanding, and intent recognition can be improved. Overall, the process of enhancing text semantics using image information can not only solve various problems existing in social media text but also provide a more rich and comprehensive perspective when predicting, analyzing, and understanding text information. This demonstrates great potential and value in multiple applications in the fields of machine learning, natural language processing, and artificial intelligence.
[0059] Although the above three methods each have their own advantages, they also have their respective deficiencies.
[0060] Speech often contains rich information, including the rhythm, pauses, and pitch changes of speech. These elements are often missing in written text, but they play an important complementary role in understanding the meaning, context, and emotions behind the text. By leveraging the pause characteristics in speech, the model can, to a certain extent, more accurately determine the boundaries of entities and improve the integrity of single-entity recognition. However, the utilization of speech information is not the key to solving all problems. Traditional text processing techniques, such as word segmentation and part-of-speech tagging, have been able to accurately divide the structure and meaning of text to a certain extent. Regarding the problem of missing context, although speech information provides a certain degree of supplementation, it does not offer a fundamental solution. Understanding context usually depends on the information in the text before and after, as well as complex factors such as other entities, sentence structures, and language usage habits in the context. In addition, the utilization of speech information also faces technical challenges. Effectively integrating speech information with text information requires the development of complex and precise speech-text collaborative processing models. These models not only need to process speech signals but also must fuse them with text data to generate richer semantic representations. This involves the cross-integration of multiple fields such as deep learning, speech recognition, and semantic understanding, and the technical implementation is difficult. Using speech information to enhance text semantic understanding is a promising method that improves the accuracy of entity recognition to a certain extent. However, this method still faces technical challenges, and its effectiveness may vary in different scenarios and tasks.
[0061] Enhancing text semantic understanding by leveraging text structure information actually guides the model to focus on the structured association information between words in the text through artificially designed features. This method emphasizes the interaction between words in terms of composition, enabling the model to better integrate the information between adjacent words and thus deepen the overall understanding of the sentence. Compared with the traditional method that only relies on the combination of text information and speech information, the utilization of text structure information is more focused on enhancing the richness and coherence within the text. However, this method of enhancing semantic understanding still faces challenges. Although it can effectively process and integrate the connections between words to a certain extent, it still struggles to handle the problem of missing context in short sentences or broken sentences. The essence of text understanding lies in understanding and parsing the context, and neither exploring the internal structure of the sentence nor combining speech information can perfectly solve the understanding problem of short sentences when context is missing.
[0062] In the combination of text information and image information, images are mainly responsible for enhancing the semantic understanding of text, especially for short phrases that depict specific scenarios or concrete things. This complementary approach can significantly improve the expressive effect of the text, enabling readers or systems to more intuitively understand the background or context described by the text, thereby reducing the difficulty and uncertainty of understanding. Images provide a visual supplement that can assist readers in constructing richer mental images and contribute to forming more vivid and comprehensive scene pictures in the brain. However, although the intuitiveness and vividness of images have their unique advantages in information transmission, there are also some limitations. The information presented by images is often more macroscopic and may involve relatively complex backgrounds and numerous elements. Although these elements enrich the background of the situation, they may also obscure the in-depth understanding of specific entities. The information provided by images is more of a "vague outline" rather than an accurate description of entities, which largely compresses the search space for the model to analyze and understand information rather than providing specific search directions or clues. From the perspective of understanding short phrases, this nature of images means that they are more in helping to construct the overall atmosphere and background of the scene and are less likely to directly reveal or explain the meaning or location of specific entities in short phrases. Therefore, although images can effectively assist text understanding, especially in enhancing the intuitiveness and contextuality of the text, for information that requires in-depth understanding of the attributes or locations of specific entities, relying solely on the description of images may be limited. To understand the text content more accurately and comprehensively, it is usually necessary to combine other forms of information (such as context, other images, or textual descriptions) to obtain more specific entity explanatory information.
[0063] To address the problem of short text named entity recognition, the present invention proposes a short text semantic enhancement method based on a domain knowledge graph and satellite images, adopting a comprehensive strategy of "short text + domain knowledge graph + satellite images".
[0064] Specifically, as a key component, the domain knowledge graph not only provides extensive background information and context associations of entities, but also captures the complex relationships and attributes between entities, thereby providing a basis for the model to gain in-depth understanding. These rich background information and relationship graphs provide strong support for entity recognition in short texts, enabling the model to more accurately identify the types and attributes of entities. On the other hand, the introduction of satellite image data enables the entire recognition system to obtain information from the visual level. Especially for those entities that are poorly described or have insufficient information in the text, through satellite images, the geographical location, shape, scale and other characteristics of the entities can be visually confirmed, further enhancing the accuracy of entity recognition. At the same time, satellite image data also provides the possibility of cross-modal fusion, that is, the combination of text and visual information, providing a new perspective and dimension for the model, and enhancing the adaptability and robustness of the system in dealing with complex scenarios. By combining short texts with domain knowledge graphs and satellite image data, the present invention not only achieves high-precision recognition of entities, but also significantly enhances the robustness and generalization ability of the system, providing innovative ideas and practical basis for the further development and application of short text named entity recognition technology.
[0065] Figure 1 The technical solution provided by the present application is shown. Generally speaking, it includes corpus annotation, model training, tuning and verification, and model use.
[0066] According to an embodiment of the present application, a short text semantic enhancement method based on a domain knowledge graph and satellite images is provided for annotating a corpus, and the corpus includes:
[0067] Short text;
[0068] Entities and types included in the short text;
[0069] Domain knowledge graph triples associated with the short text;
[0070] Satellite images associated with the short text.
[0071] The semantic enhancement method includes:
[0072] S1: Use a large language model to assist in annotating the entities and types in the short text to improve the annotation speed;
[0073] S2: Adopt multi-way recall and re-ranking means to retrieve domain knowledge graph triples associated with the short text from the domain knowledge graph, and use the sentence-formed triples as entity knowledge in short text entity recognition. As Figure 2 shown, specifically including:
[0074] S21: Preprocess the original short text, and the preprocessing includes keyword extraction and text vectorization to obtain keywords and short text vectors;
[0075] S22: Recall a series of candidate subgraphs from the domain knowledge graph and retrieve the set of triples associated with the short text;
[0076] S23: Vectorize the triples in the triple set using a short text vectorization model, calculate the cosine distance between each triple and the short text vector as the semantic relevance score, sort them in descending order of the semantic relevance score, select the top N triples (i.e., the N triples with the highest scores), sentence-ize the N triples with the highest scores, and use the sentence-ized N triples as entity knowledge in short text entity recognition.
[0077] S3: Determine whether the short text in the original corpus contains attached satellite images: If the short text in the original corpus has attached satellite images, directly use the attached satellite images; otherwise, search for satellite images associated with the short text in the satellite image library and use them as background knowledge for short text entity recognition, as Figure 3 shown below, specifically including:
[0078] S31: Tokenize and vectorize the short text;
[0079] S32: Conduct multi-way retrieval in the satellite image library to recall several satellite images related to the short text;
[0080] S33: Convert the satellite image into text using a judgment model and perform vectorization processing on the text to obtain a satellite image vector, and then calculate the cosine distance between the satellite image vector and the short text vector as the correlation score between the satellite image and the short text;
[0081] S34: Sort the scores from high to low and select the satellite image with the highest score as background knowledge for short text entity recognition.
[0082] S4: Sentence-ize the triples and satellite images respectively to unify the input form and obtain the text with enhanced semantics.
[0083] The sentence-ization rule for triples is: "{Entity name}'s {Attribute / relationship name} is {Attribute value / relationship value}", for example: the triple "Adora Cruises - Belonging company - CSSC Carnival Cruise Co., Ltd." is converted to "The belonging company of Adora Cruises is CSSC Carnival Cruise Co., Ltd.".
[0084] The sentence-ization rule for satellite images is: Use models such as image interpretation and object detection to extract information such as location, space, target, layout, and environment from the satellite image, and summarize the information as the background information of the short text.
[0085] S5: Use entity knowledge and background knowledge to perform entity annotation on the semantically enhanced text, give real entity labels, and obtain the annotated corpus.
[0086] Add prefixes respectively to facilitate better understanding by the model. For example, an original short text "The Adora Magic City leaves Qingdao" becomes "Short sentence: The Adora Magic City leaves Qingdao" after processing. <s>Entity knowledge: The company to which Adora Cruises belongs is CSSC Carnival Cruise Lines Co., Ltd.\nBackground: Near Qingdao Port, there is a sea in the distance and a cruise ship nearby, etc. The corresponding true entity labels are: "Cruise ship: Adora Magic City, Port: Qingdao, Movement: Departing". Among them, " <s>", "\n" is used as a delimiter, and "short sentence", "entity knowledge", and "background" are used as semantic prefixes to help the model understand the semantics and functions of different segments.
[0087] According to an embodiment of the present application, a short text understanding model is also provided. The model architecture includes BERT (Bidirectional Encoder Representations from Transformers), BI-LSTM (Bidirectional Long Short-Term Memory), and CRF (Conditional Random Field), so as to construct an efficient and accurate short text entity recognition system.
[0088] Among them, BERT is used for feature extraction and fusion, and is trained based on large-scale unlabeled text. BERT can capture rich context information and semantic features, providing a strong semantic representation for subsequent processing.
[0089] Among them, BI-LSTM can effectively improve the model's understanding and prediction ability for sequence tasks by considering the forward and backward dependencies of sequence data.
[0090] Finally, CRF has a powerful sequence labeling ability, which can ensure that the model output not only considers the optimality of the current state, but also takes into account the global optimization of the entire sequence, so as to achieve better prediction performance when outputting sequence labels.
[0091] According to an embodiment of the present application, a training method for a short text understanding model is also provided. The labeled corpus obtained by using the above semantic enhancement method is used as the training corpus for model training. The training method includes:
[0092] S1: Split the labeled corpus into a training set, a test set, and a validation set, and the ratio is set to 3:1:1. The training set is used for model training, the test set is used to verify the generalization ability of the model, and the validation set is used in the model adjustment stage to select the optimal model parameters and avoid overfitting.
[0093] S2: Before model training, set the key hyperparameters of the model, including the learning rate, regularization coefficient, etc. The reasonable setting of these hyperparameters is crucial for the model performance.
[0094] S3: Start iterative training to gradually optimize the short text understanding model under the guidance of these hyperparameters, and finally reach the best performance state.
[0095] The model training follows the data splitting strategy to ensure the fairness and effectiveness of the training process.
[0096] According to an embodiment of the present application, a usage method for a short text understanding model is also provided. After the model training is completed, it can be used in a production environment. Before use, preprocessing is performed in the same way as when labeling data, including:
[0097] Retrieve triples related to short texts in the domain knowledge graph and provide satellite images related to the short texts;
[0098] Semanticize the retrieved triples and images, that is, organize the triples into phrases, such as "The company to which the Adora Cruises belongs is CSSC Carnival Cruise Shipping Limited", and the images are processed through image interpretation and object detection models to extract information such as location, space, and objects, which are supplemented as background information behind;
[0099] Input the processed sentences into the model to obtain the entities and corresponding types that appear in the short text.
[0100] The above are only illustrative specific embodiments of the present application and are not intended to limit the scope of the present application. Any equivalent changes, modifications, and combinations made by those skilled in the art without departing from the concept and principles of the present application shall fall within the scope of protection of the present application.< / s> < / s>
Claims
1. A short text semantic enhancement method based on domain knowledge graph and satellite imagery, comprising: S1: Use a large language model to assist in labeling entities and types in short texts; S2: Use multi-way recall and rearrangement to retrieve domain knowledge graph triples associated with short texts from the domain knowledge graph, and use the sentence-based triples as entity knowledge in short text entity recognition; S3: Determine whether the short text of the original corpus contains satellite images: If the short text of the original corpus contains satellite images, the satellite images are used directly; otherwise, the satellite images associated with the short text are searched in the satellite image library as background knowledge for short text entity recognition; S4: process the triples and satellite images into sentences respectively to unify the input form and obtain the semantically enhanced text; S5: Use entity knowledge and background knowledge to perform entity annotation on the semantically enhanced text, give real entity labels, and obtain annotated corpus.
2. The method according to claim 1, wherein step S2 comprises: S21: preprocessing the original short text, including keyword extraction and text vectorization, to obtain keyword and short text vectors; S22: Recall a series of candidate subgraphs from the domain knowledge graph in multiple ways and retrieve the set of triples associated with the short text; S23: Use the short text vectorization model to vectorize the triples in the triple set, calculate the cosine distance between each triple and the short text vector as the semantic relevance score, sort them from high to low according to the semantic relevance score, select the N triples with the highest scores, convert the N triples with the highest scores into sentences, and use the N triples after sentence conversion as entity knowledge in short text entity recognition.
3. The method according to claim 1, wherein step S3 comprises: S31: Segment and vectorize short texts; S32: Perform multi-channel retrieval in the satellite image library to recall several satellite images related to the short text; S33: converting the satellite image usage judgment model into text and performing vectorization processing on the text to obtain a satellite image vector, and then calculating the cosine distance between the satellite image vector and the short text vector as a correlation score between the satellite image and the short text; S34: Sort the scores from high to low, and select the satellite image with the highest score to be used as background knowledge for short text entity recognition.
4. The method according to claim 1, used to annotate a corpus, the corpus comprising: Short text; Entities and types contained in the short text; Domain knowledge graph triples associated with short texts; Satellite imagery associated with a short text.
5. The method according to claim 1, wherein in step S4, The sentence processing rule of triples is: {attribute / relation name} of {entity name} is {attribute value / relation value}; The sentence processing rules of satellite images are as follows: use image interpretation, target detection and other models to extract location, space, target, layout, environment and other information from satellite images, and summarize the information as background information of short texts.
6. A short text understanding model, the model architecture of which includes: Bidirectional pre-trained language model for feature extraction and fusion, based on large-scale unlabeled text training, captures contextual information and semantic features, and provides semantic representation for subsequent processing; Bidirectional long short-term memory network, used to consider the dependencies between sequence data; Conditional Random Fields.
7. A training method for a short text understanding model, using the annotated corpus obtained by the semantic enhancement method of claim 1 as training corpus for model training, the training method comprising: S1: Split the annotated corpus into training set, test set, and validation set; S2: Set key hyperparameters of the model; S3: Start iterative training to gradually optimize the short text understanding model under the guidance of hyperparameters to achieve the best performance.
8. The method according to claim 7, wherein: The ratio of training set, test set, and validation set is set to 3:1:
1.
9. The method according to claim 7, wherein: The training set is used to train the model, the test set is used to verify the generalization ability of the model, and the validation set is used to select the optimal model parameters during the model adjustment stage.
10. The method according to claim 7, wherein: The key hyperparameters of the model include learning rate and regularization coefficient.
Citation Information
Cited By
Chinese semantic matching enhancement method and system based on BAAI-bge model
CN120578756A