A semantic drift detection method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202310224345.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-03-09
AI Technical Summary
但它的假设过于肯定,因此会引入大量的噪声,存在语义漂移现象
[0018]本发明实施例的技术方案,通过获取待识别文本数据,根据预设知识抽取框架获取待识别文本数据中实体文本以及实体关系,基于预设语义漂移检测模型对实体语义以及实体关系进行语义漂移检测,确定语义漂移情况,实现检测电力领域数据的语义漂移情况,降低人工检测的成本,进而可以剔除低质量的数据,构建高质量电力领域知识图谱。
Smart Images

Figure CN116502646B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a semantic drift detection method, apparatus, electronic device, and storage medium. Background Technology
[0002] The knowledge graph in the power sector aims to fully utilize the data information carried by the Internet of Things in the power sector to characterize the concepts, entities, events and their relationships in the power system in a structured way, providing the power industry with a more effective cross-media big data organization, management and cognition capability.
[0003] In constructing a knowledge graph, knowledge extraction is required from data of different sources and structures to form structured data that is then stored within the knowledge graph. To reduce reliance on manually labeled data, knowledge bases can be aligned with unstructured text to automatically build large amounts of training data. However, this approach relies on overly certain assumptions, introducing significant noise and semantic drift. When semantic drift exists in power sector data, it can lead to inaccurate data in the construction of power sector knowledge graphs, resulting in weak correlations within the knowledge graph. This can pose safety risks to power industry personnel using power sector knowledge graphs in their work. Therefore, detecting semantic drift in power sector data, eliminating low-quality data, and constructing high-quality power sector knowledge graphs have become urgent problems to be solved. Summary of the Invention
[0004] This invention provides a semantic drift detection method, device, electronic device, and storage medium to achieve rapid detection of semantic drift in power data, facilitating the construction of high-quality power knowledge graphs.
[0005] According to one aspect of the present invention, a semantic drift detection method is provided, wherein the method includes:
[0006] Obtain the text data to be recognized;
[0007] The entity types and entity relationships of the entity text in the text data to be identified are obtained according to the preset knowledge extraction framework. The knowledge extraction framework includes an entity extraction framework and an entity relationship extraction framework.
[0008] The semantic drift detection model is used to detect semantic drift in entity types and entity relationships to determine the semantic drift situation. The preset semantic drift detection model is generated by training on labeled power datasets, power seed sets and unlabeled power data.
[0009] According to another aspect of the present invention, a semantic drift detection device is provided, characterized in that it comprises:
[0010] The text data acquisition module is used to acquire the text data to be recognized;
[0011] The entity acquisition module is used to acquire the entity type and entity relationship of the entity text in the text data to be identified according to the preset knowledge extraction framework. The knowledge extraction framework includes an entity extraction framework and an entity relationship extraction framework.
[0012] The semantic drift detection module is used to detect semantic drift of entity types and entity relationships based on a preset semantic drift detection model, and to determine the semantic drift situation. The preset semantic drift detection model is trained and generated based on labeled power datasets, power seed sets and unlabeled power data.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] At least one processor;
[0015] and memory that is communicatively connected to at least one processor;
[0016] The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by at least one processor so that at least one processor can execute the semantic drift detection method of any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the semantic drift detection method of any embodiment of the present invention.
[0018] The technical solution of this invention obtains text data to be identified, acquires entity text and entity relationships in the text data to be identified according to a preset knowledge extraction framework, performs semantic drift detection on entity semantics and entity relationships based on a preset semantic drift detection model, determines the semantic drift situation, realizes the detection of semantic drift in power field data, reduces the cost of manual detection, and can then eliminate low-quality data and construct a high-quality power field knowledge graph.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a semantic drift detection method provided in Embodiment 1 of the present invention;
[0022] Figure 2 This is a flowchart of a semantic drift detection method provided in Embodiment 2 of the present invention;
[0023] Figure 3 This is a flowchart of the training process for a preset semantic drift detection model provided in Embodiment 3 of the present invention;
[0024] Figure 4 This is a flowchart of the training process of an entity semantic detection model according to Embodiment 3 of the present invention;
[0025] Figure 5 This is a flowchart of the training process for a relational semantic detection model provided in Embodiment 3 of the present invention;
[0026] Figure 6 This is an architecture diagram for semantic drift detection provided according to Embodiment 4 of the present invention;
[0027] Figure 7 This is a schematic diagram of a preset knowledge extraction framework provided in Embodiment 4 of the present invention;
[0028] Figure 8 This is a schematic diagram of the structure of a preset semantic drift detection model provided in Embodiment 4 of the present invention;
[0029] Figure 9 This is a schematic diagram of the Transformer encoder provided in Embodiment 4 of the present invention;
[0030] Figure 10 This is a schematic diagram of the structure of a Block according to Embodiment 4 of the present invention;
[0031] Figure 11 This is a schematic diagram of the structure of a semantic drift detection device provided in Embodiment 5 of the present invention;
[0032] Figure 12 This is a schematic diagram of the structure of an electronic device that implements the semantic drift detection method of this invention. Detailed Implementation
[0033] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0035] Example 1
[0036] Figure 1 This is a flowchart of a semantic drift detection method according to Embodiment 1 of the present invention. This embodiment is applicable to detecting semantic drift in data text. The method can be executed by a semantic drift detection device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0037] S110. Obtain the text data to be recognized.
[0038] The text data to be identified can refer to text data awaiting semantic drift detection. In practice, the text data to be identified can include structured data, semi-structured data, and unstructured data in the power sector. For example, professional knowledge in the power sector and the models of various power equipment are all considered text data to be identified. Of course, the text data to be identified is not limited to the power sector; data from the financial and consumer sectors are also within the scope of this invention. The text to be identified can contain one or more entity texts, and there can be certain relationships between different entity texts.
[0039] In the embodiments of the invention, the text data to be identified can be stored locally on the electronic device or in a database, and can be retrieved from the local device or database. In actual operation, the file storing the text data to be identified can be selected on the local device to obtain the text data; alternatively, the text data stored in the database can be extracted as the text data to be identified; or, power data can be downloaded from a power data website as the text data to be identified.
[0040] S120. Obtain the entity type and entity relationship of the entity text in the text data to be identified according to the preset knowledge extraction framework, wherein the knowledge extraction framework includes an entity extraction framework and an entity relationship extraction framework.
[0041] The preset knowledge extraction framework can be a pre-set extraction framework for extracting knowledge from data of different sources and structures. The preset knowledge extraction framework can include an entity extraction framework and an entity relationship extraction framework. The entity extraction framework can extract entity text and corresponding entity types from the text data to be identified, while the entity relationship extraction framework can extract entity relationships from the text data to be identified. In other words, the preset extraction framework can simultaneously obtain entity text and entity relationships from the text data to be identified. Entity text can refer to text with entity meaning, and entity types can be used to describe entity features or characteristics. In one embodiment, entity text can include, but is not limited to, specific electrical equipment names, dates, and times. For example, entity text can include, but is not limited to, transformers, equipment models, etc.; entity types can include, but are not limited to, equipment, locations, facilities, etc.; entity relationships can refer to the semantic relationships between entities in the text data, and can include, but are not limited to, physical location relationships, part-whole relationships, and subordinate relationships.
[0042] In the embodiments of the invention, a preset knowledge extraction framework can be extracted to obtain entity texts and entity relationships in the text data to be identified. In actual operation, a piece of text data to be identified may include one or more entity texts and entity relationships. Each entity text corresponds to an entity type. The text data to be identified is input into the preset knowledge framework to extract the entity texts and entity relationships contained in the text data to be identified, thus determining the entity type of the entity texts. Entity texts can be extracted from the text data to be identified using an entity extraction framework, and entity relationships can be extracted from the text data to be identified using an entity relationship framework. The preset knowledge extraction framework can be composed of a neural network or a named entity recognition tool; for example, it may include a feedforward neural network, Label Studio, the Python jieba library, etc.
[0043] In one embodiment, when the preset knowledge extraction framework is based on a neural network, the text data to be identified can be serialized, and part-of-speech tagging can be performed on the text data to be identified using a bidirectional long short-term memory network. The probability of each phrase tag type is calculated based on the softmax loss function of the feedforward neural network to determine the entity text and its corresponding entity type. In one embodiment, BIO tagging or BIOES tagging can also be used to determine the entity text in the text data to be identified. In actual operation, after part-of-speech tagging of the text data to be identified, a classifier can be used to determine the character positions of entities, and the nearest matching principle can be used to pair the relationships between entities to determine the entity relationships in the text data to be identified. In one embodiment, the preset knowledge extraction framework can generate triplet tagging data of entity text 1, entity text 2, and the entity relationships between the two entity texts. In one embodiment, when the preset knowledge extraction framework is a named entity tagging tool such as Label Studio, the text data to be identified can be input into the named entity tagging tool, and the named entity tagging tool can determine the entity types and entity relationships of the entity texts in the text data to be identified.
[0044] S130. Based on the preset semantic drift detection model, perform semantic drift detection on entity types and entity relationships to determine the semantic drift situation. The preset semantic drift detection model is generated by training on labeled power datasets, power seed sets and unlabeled power data.
[0045] The pre-defined semantic drift detection model refers to a pre-set model for detecting semantic drift. This model can be trained and generated based on labeled electricity datasets, electricity seed datasets, and unlabeled electricity data. In practice, the pre-defined semantic drift detection model can be built based on a Transformer network. This model can determine whether semantic drift has occurred in entity semantics and entity relationships. The labeled electricity dataset can include pre-labeled electricity data pairs, such as triplet-labeled data, which can be used as the training set to train the pre-defined semantic drift detection model. The unlabeled electricity data can be labeled using the electricity seed set to serve as the test set for testing the training of the pre-defined semantic drift detection model. Semantic drift can be categorized as either semantic drifted or not.
[0046] In practice, a pre-created semantic drift detection model can be extracted. The extracted entity text, entity type, and entity relationship are input into the pre-created semantic drift detection model to determine whether semantic drift has occurred in the entity type and entity relationship. The pre-created semantic drift detection model can be generated by training on labeled electricity datasets, electricity seed sets, and unlabeled electricity data. In practice, semantic drift is considered to have occurred when either the entity type or entity relationship drifts; conversely, semantic drift is considered to have not occurred when neither the entity type nor the entity relationship drifts.
[0047] In this embodiment of the invention, by acquiring text data to be identified, entity texts and entity relationships in the text data to be identified are obtained according to a preset knowledge extraction framework. Based on a preset semantic drift detection model, semantic drift detection is performed on entity semantics and entity relationships to determine the semantic drift situation. This enables convenient detection of semantic drift in power field data, reduces the cost of manual detection, and can eliminate low-quality data to construct a high-quality knowledge graph in the power field.
[0048] In one embodiment, after obtaining the entity text and entity relationships in the text data to be identified according to a preset knowledge extraction framework, the method further includes: performing entity alignment on the entity text.
[0049] Entity alignment can determine whether two or more entities from different information sources refer to the same object in the real world, and can group entity texts with the same name together. In practice, entity text extracted through a pre-set knowledge framework may be incomplete; entity alignment can be used to obtain complete entity text.
[0050] In the embodiments of the invention, entity alignment methods can include various methods, including but not limited to Elasticsearch full-text search. In actual operation, entity text can be segmented into words, and the segmented entity text can be used as keywords to perform a full-text search in the text data to be identified, obtaining text containing entity text as candidate text. The entity text is then decomposed into multiple segments, and similarity is calculated between each segment and the candidate entities. The candidate entities are then ranked, and the entity with the highest output score is taken as the aligned entity text.
[0051] Example 2
[0052] Figure 2 This is a flowchart of a semantic drift detection method according to Embodiment 2 of the present invention. This embodiment is a further description of a semantic drift detection method based on the above embodiments. Figure 2 As shown, the method includes:
[0053] S210. Obtain the text data to be recognized.
[0054] S220. The text data to be identified is serialized to generate a text sequence to be identified, and the bidirectional long short-term memory network in the preset knowledge extraction framework is called to perform part-of-speech tagging on the text sequence to be identified.
[0055] The text sequence to be identified can refer to the sequence of text and numbers generated by text serialization of the text data to be identified. In practice, the method of text serialization is not limited. The preset knowledge extraction framework can be pre-set and is used to extract the text data to be identified. In practice, the preset knowledge extraction framework can include an entity extraction framework and an entity relation extraction framework. The bidirectional long short-term memory network consists of a forward long short-term memory network and a backward long short-term memory network.
[0056] In this embodiment, the text data to be recognized can be processed through text serialization to generate a text sequence to be recognized. A bidirectional long short-term memory (BSSM) network within a preset knowledge extraction framework is then invoked, and the encoder of the BSSM network annotates the character positions of the text sequence to be recognized. In actual operation, the text sequence to be recognized can be input into the BSSM network, and the encoder of the BSSM network can perform part-of-speech tagging at each position in the text sequence to be recognized. In one embodiment, the BIO tagging method or the BIOES tagging method can also be used to determine the tagging of each position in the text data to be recognized.
[0057] S230. Determine the entity text of the text data to be identified and the entity type of the corresponding entity text according to the softmax loss function of the feedforward neural network in the preset knowledge extraction framework.
[0058] In this embodiment, the entity text and its corresponding entity type can be determined by invoking a feedforward neural network within the knowledge extraction framework, based on a softmax loss function. In practice, the part-of-speech-tagged text data is input into the feedforward neural network. The feedforward neural network calculates the tagging type probability for each word based on the softmax loss function, thus determining the entity text and its corresponding entity type. In one embodiment, the tagging type corresponding to the highest tagging type probability can be used as the entity text and its corresponding entity type.
[0059] S240. Call the prediction classifier based on the feedforward neural network in the preset knowledge extraction framework to classify the part-of-speech tags and determine the relational semantics corresponding to the entity text in the text data to be identified.
[0060] The predictive classifier can be a classifier based on a feedforward neural network, which can be used to determine whether the part-of-speech tagging result at each position is the start or end position of the entity text.
[0061] In one embodiment, a predictive classifier based on a feedforward neural network within a pre-defined knowledge extraction framework can be invoked to determine whether the part-of-speech tagging result at each position corresponds to the start or end position of the entity text. After identifying the entity text, a proximity matching principle can be used to pair the entity texts, predicting the relationship between the two entity texts based on the pre-defined knowledge extraction framework, thus determining the relational semantics corresponding to the entity text. In one embodiment, the relationship between two entity texts can be predicted based on the softmax loss function of the feedforward neural network. The probability of each semantic relationship is determined using the softmax loss function, and the semantic relationship corresponding to the maximum probability value is taken as the relational semantics corresponding to that entity text, thereby determining the relational semantics corresponding to each entity text in the text data to be identified. In one embodiment, before inputting to the predictive classifier based on the feedforward neural network, stacking can be performed to allow different networks to extract different information from the data.
[0062] S250. Align the entity text with entities.
[0063] S260. Based on the preset semantic drift detection model, perform semantic drift detection on entity semantics and entity relationships to determine the semantic drift situation.
[0064] In the embodiments of the invention, a pre-created preset semantic drift detection model can be extracted, and the entity text, entity type and entity relationship after entity alignment can be input into the preset semantic drift detection model to determine whether semantic drift occurs in the entity type and entity relationship.
[0065] In one embodiment, semantic drift can include at least one of the following:
[0066] When the similarity value between the entity type in the text data to be identified and the preset entity type is greater than the preset similarity threshold, it is confirmed that the entity type has not undergone semantic drift.
[0067] When the entity relationship output by the preset semantic drift detection model is contained in the text data to be identified, it is confirmed that the entity relationship has not undergone semantic drift.
[0068] If neither the entity type nor the entity relationship has undergone semantic shift, it is confirmed that the text data to be identified has not undergone semantic shift.
[0069] When both entity type and / or entity relationship undergo semantic shift, it is confirmed that the text data to be identified has undergone semantic shift.
[0070] The preset entity type can refer to pre-stored, already identified entity types. The preset similarity threshold can be a parameter used to determine whether semantic drift has occurred in the entity types of the text data to be identified. When the similarity value between an entity type and the preset entity type is greater than the preset similarity threshold, it can be considered that the entity type has not undergone semantic drift; when the similarity value between an entity type and the preset entity type is less than the preset similarity threshold, it can be considered that the entity type has undergone semantic drift. The preset similarity threshold can be pre-set or set by the detection personnel according to the semantic drift detection requirements. The higher the preset similarity threshold, the higher the accuracy of semantic drift detection.
[0071] In this embodiment, after entity text, entity type, and entity semantics are input into a preset semantic drift detection model, the entity text can be converted into corresponding entity vectors and semantic vectors through an embedding layer. Based on the entity text, a pre-stored preset entity text is queried, and the similarity value between the entity text and the extracted preset entity text is calculated. When the similarity value between the entity type and the preset entity type is greater than a preset similarity threshold, the entity type is considered to be the same as the pre-stored preset entity type, confirming that no semantic drift has occurred. The semantic vector can be processed by a fully connected network for multi-classification, outputting probability values for each classification task. The value corresponding to the maximum value is selected as the predicted label, and the predicted label is used as the entity relation output by the preset semantic drift detection model. When the output entity relation is included in the text data to be identified, it can be confirmed that no semantic drift has occurred in the entity relation. In actual operation, when neither entity type nor entity relationship has undergone semantic drift, the text data to be identified can be considered to have undergone semantic drift; when either entity type or entity relationship has undergone semantic drift, the text data to be identified can be considered to have undergone semantic drift; when both entity type and entity relationship have undergone semantic drift, the text data to be identified can be considered to have undergone semantic drift.
[0072] This invention, in its embodiments, acquires text data to be identified, serializes the text data to generate a text sequence to be identified, calls a bidirectional long short-term memory network within a preset knowledge extraction framework to perform part-of-speech tagging on the text sequence, identifies and determines the entity text and corresponding entity types of the text according to the softmax loss function of the feedforward neural network within the preset knowledge extraction framework, calls a predictive classifier based on the feedforward neural network within the preset knowledge extraction framework to classify the part-of-speech tags, determines the relational semantics corresponding to the entity text in the identified text, aligns the entity text, and performs semantic drift detection on the entity semantics and entity relationships based on a preset semantic drift detection model to determine the semantic drift situation. This enables the joint extraction of entity text and entity relationships through the preset knowledge extraction framework, identifying multiple potential relationships for each entity in the power sector, thus improving knowledge extraction capabilities. The semantic drift detection model, by performing semantic drift detection on the entity semantics and entity relationships, more quickly determines the semantic drift situation of power sector data.
[0073] Example 3
[0074] Figure 3 This is a flowchart illustrating the training process of a preset semantic drift detection model according to Embodiment 3 of the present invention. This embodiment is applicable to training a preset semantic drift detection model, which includes an entity semantic detection model and a relation semantic detection model. The preset semantic drift detection model includes an input layer, a shared layer, and two task layers, as follows: Figure 3 As shown, the training of the preset semantic drift detection model includes:
[0075] S310. Obtain the pre-stored labeled power dataset and use it as the training set.
[0076] The labeled power dataset can be a pre-constructed dataset used to store labeled entity types and entity types in the power domain. In one embodiment, the data stored in the labeled power dataset may include labeled power domain text data, wherein the labeled power domain text data may include labeled entity text, the entity type corresponding to the entity text, whether any two entities have the same entity relationship, and the entity relationship. In one embodiment, whether any two entities have the same entity relationship can be marked by similarity tags. For example, when two entity texts have the same entity relationship, the similarity tag can be 1; when two entity texts have different entity relationships, the similarity tag can be 0.
[0077] In the embodiments of the invention, the labeled power dataset can be pre-stored on the local device or cloud server. The labeled power dataset can be searched on the local device or cloud server and extracted as a training set for training a preset semantic drift detection model.
[0078] S320. Input the training set into the pre-constructed preset semantic drift detection model for training, wherein the comprehensive loss function of the preset semantic drift detection model is determined by the entity semantic detection model and the relation semantic detection model.
[0079] The preset semantic drift detection model includes an entity semantic detection model and a relational semantic detection model. The preset semantic drift detection model includes an input layer, a shared layer, and two task layers.
[0080] The preset semantic drift detection model is a multi-task model. The entity semantic detection model can be used to detect whether semantic drift has occurred in entity semantics; the relation semantic detection model can be used to detect whether semantic drift has occurred in relation semantics. The preset semantic monitoring model can be composed of the entity semantic detection model and the relation semantic detection model. The preset semantic drift detection model can include an input layer, a shared layer, and two task layers. In actual operation, the input layer can be used to process the input entity text, entity type, and entity relation. The shared layer can be used to vectorize the entity text for feature extraction in subsequent task layers. The task layers can include an entity semantic detection task layer and a relation semantic detection task layer, where the entity semantic detection task layer can be used to detect semantic drift in entity semantics, and the relation semantic detection task layer can be used to detect semantic drift in relation semantics. In one embodiment, the preset semantic drift detection model can be built based on a Transformer network, and the shared layer includes an Embedding layer and a TransformerEncoder layer. The entity semantic detection task layer can calculate the similarity value of two entity texts to determine the semantic drift of entities. The relational semantic detection task layer can redetermine entity relationships through a fully connected network, thereby determining semantic drift.
[0081] The comprehensive loss function can be a computational function used to measure the difference between the predicted value and the true value of the preset semantic drift detection model. The smaller the loss function, the better the robustness of the model. The comprehensive loss function of the preset semantic drift detection model can be determined by the entity semantic detection model and the relation semantic detection model. In one embodiment, the comprehensive loss function can be determined by the entity semantic detection model and the relation semantic detection model. For example, the comprehensive loss function may include:
[0082]
[0083] Where σ1 and σ2 are noise parameters, which control the relative weights of the L1(W) and L2(W) losses respectively. If the noise parameter σ is larger, the weight of the corresponding loss function L(W) will be smaller. However, since the model will try to make the loss function as small as possible, σ will become very large, completely ignoring the influence of the data. Therefore, a regularization term logσ is added to the noise term.
[0084] In the embodiments of the invention, the training set can be input into a pre-built preset semantic drift detection model to train the pre-built preset semantic drift detection model until the preset value of the comprehensive loss function is reached, thus completing the training of the preset semantic drift detection model.
[0085] In this embodiment of the invention, a pre-stored labeled power dataset is obtained, which is then used as a training set. The training set is input into a pre-built preset semantic drift detection model for training. This completes the training of the pre-built preset semantic drift detection model, which can then detect semantic drift and improve the user experience.
[0086] In one embodiment, the training of the preset semantic drift detection model includes the training of an entity semantic detection model and the training of a relational semantic detection model.
[0087] In one embodiment, Figure 4 This is a flowchart of the training process for an entity semantic detection model according to Embodiment 3 of the present invention, as follows: Figure 4 As shown, the training of the entity semantic detection model includes:
[0088] S321. Input the two entity texts and similar labels from the training set into the entity semantic detection model to obtain the entity vectors corresponding to the training set.
[0089] The entity semantic detection model uses a Siamese network architecture, where the two entity texts can be any two entity texts from the same data in the training set. Similarity labels are determined based on the degree of similarity between the entity types of the two entity texts, and can include both similar and dissimilar labels. In practice, similarity labels can be represented by real numbers in the range [0,1], where similarity can be 1 and dissimilarity can be 0. A larger value indicates a more similar entity type between the two entity texts. Entity vectors can refer to the vectors corresponding to the entity texts; different entity texts can correspond to different entity vectors. In one embodiment, the input format for the two entity texts and similarity labels can be two entity texts and a similarity label, separated by a \t delimiter.
[0090] In this embodiment, two entity texts from the same data in the training set can form an instance pair, and each instance pair can correspond to a similar label. The two entity texts and their similar labels from the training set can be input into the entity semantic detection model to determine the entity vectors corresponding to the training set. In actual operation, after the entity texts are input into the entity semantic detection model, the positive integer indices of the entity texts in the source data can be determined first. These positive integer indices are then transformed using one-hot vectors to capture the relationships between texts and determine the initial text vectors. The encoder in the Transformer then adds the residual to determine the entity vector corresponding to each entity text in the training set. The encoder in the Transformer can be considered as composed of multiple blocks, each block adding residual connections, Layer Norm, and Full-Connection (FC) on top of Self-attention, to more accurately determine the entity vectors corresponding to the training set.
[0091] S322. Call the preset function to determine the similarity value between the entity vectors of different entity texts.
[0092] The preset function can be a pre-defined function used to determine the similarity value between entity vectors of different entities. In actual operation, the preset function can include, but is not limited to, the distance function and the cosine function. The similarity value can be a real number with a value range of [0,1], indicating the probability that two entity texts share the same relation type. The higher the score, the greater the probability that the two instances express the same relation.
[0093] In the embodiments of the invention, a preset function can be extracted, and the similarity value between entity vectors of different entities can be determined based on the preset function. In actual operation, when the preset function is a distance function, the distance function may include s(x,y)=σ(w s T (f s (x)-f s (y)) 2 +b s In this function, fs(x) and fs(y) represent the output functions of the encoder, and x and y represent entity vectors. σ() represents the sigmoid function, ws represents the weights, and bs represents the biases. The similarity between the entity vectors of two entity texts is calculated by inputting them into the preset function.
[0094] S323. When the similarity value is greater than the preset similarity threshold, determine that the entity semantic detection model training is complete; otherwise, determine the average absolute error loss between the similarity value and the preset similarity threshold.
[0095] The preset similarity threshold can be a pre-set threshold used to determine whether the entity semantic detection model has completed training. The preset similarity threshold can be determined according to the needs of business personnel.
[0096] In this embodiment of the invention, when the similarity value is greater than a preset similarity threshold, the entity semantic detection model can be considered to have completed training, and training of the entity semantic detection model can be stopped at this time. When the similarity value is less than or equal to the preset similarity threshold, the entity semantic detection model can be considered not to have completed training, and the average absolute error loss between the similarity value and the preset similarity threshold can be calculated. The average absolute error loss can be used as the loss function of the entity semantic detection model. For example, the average error loss function may include: Where h(x) represents the predicted score and y represents the true score (0,1).
[0097] S324. After optimizing the weights and parameters of the entity semantic detection model according to the mean absolute error loss, retrain the entity semantic detection model.
[0098] In the embodiments of the invention, the weights and parameters of the entity semantic detection model can be optimized according to the mean absolute error loss. After optimizing the weights and parameters, the entity semantic detection model can be retrained according to the above steps until the similarity value is greater than the preset similarity threshold, thus completing the training of the entity semantic detection model.
[0099] In this embodiment of the invention, by inputting two entity texts and similar labels from the training set into the entity semantic detection model, the entity vectors corresponding to the training set are obtained. A preset function is called to determine the similarity value between the entity vectors of different entity texts. When the similarity value is greater than a preset similarity threshold, the entity semantic detection model is considered to have completed training. Otherwise, the mean absolute error loss between the similarity value and the preset similarity threshold is determined. The weights and parameters of the entity semantic detection model are optimized according to the mean absolute error loss, and the entity semantic detection model is retrained. This achieves the training of the entity semantic detection model. By using the mean absolute error loss function as the objective function for optimization, the accuracy of the entity semantic detection model is improved, enhancing the user experience.
[0100] In one embodiment, Figure 5 This is a flowchart of the training process for a relational semantic detection model according to Embodiment 3 of the present invention, as follows: Figure 5 As shown, the training of the relation semantic detection model includes:
[0101] S325. Input the entity relations and relation labels in the training set into the relation semantic detection model to obtain the semantic vector corresponding to the training set.
[0102] Here, the relation vector can refer to the vector corresponding to entity relations, and different entity relations can correspond to different semantic vectors. The relation label can include positive sample labels and negative sample labels. By training with negative samples, the false detection rate and false recognition rate can be reduced, and the generalization ability of the network model can be improved. In one embodiment, the input format for entity relations and relation labels can be entity relations and relation labels, separated by the \t delimiter.
[0103] In this embodiment of the invention, after the entity relation is input into the entity semantic detection model, the positive integer index of the text corresponding to the entity relation in the source data can be determined first. This positive integer index is then transformed using a one-hot vector to capture the relationship between texts and determine the initial semantic vector. Finally, the encoder in the Transformer adds this vector to the residual to determine the semantic vector corresponding to each entity relation text in the training set.
[0104] In one embodiment, during actual operation, S321 and S325 can simultaneously input two entity texts, similar labels, entity relationships, and relationship labels from the training set into a preset semantic drift detection model to determine the entity vector and semantic vector corresponding to the training set.
[0105] S326. Call the fully connected network to determine the multi-classification of different semantic vectors, generate at least two classification task probability values, and select the label with the maximum value as the predicted label.
[0106] The fully connected network is the most basic layer in a neural network / deep neural network, where each node is connected to all nodes in the layer above. Fully connected networks can be used for multi-classification of different semantic vectors.
[0107] In this embodiment, a semantic vector can be input into a fully connected network. The fully connected network performs multi-classification on the semantic vector, determines the probability value for each classification task, and determines the predicted label. In actual operation, after the semantic vector is input into the fully connected network, the network can classify the semantic vector, evaluate the predicted value for each classification task, and select the label with the maximum value as the predicted label.
[0108] S327. When the correct probability value of the predicted label is greater than the preset probability threshold, determine that the entity semantic detection model training is complete; otherwise, determine the cross-entropy loss between the correct probability value and the preset probability threshold.
[0109] The preset probability threshold can be pre-set and used to determine whether the relation semantic detection model has completed training. This threshold can be determined based on the needs of business personnel. The correct probability value can be determined by comparing the predicted label with the input relation label. If the predicted label matches the relation label, the predicted label is considered correct. The correct probability value of the predicted label can be determined by dividing the number of correct predicted labels by the total number of predicted labels.
[0110] In this embodiment, the correct probability value can be determined based on the predicted label and the input relationship label. When the correct probability value is greater than a preset probability threshold, the relationship semantic detection model is considered to have completed training, and training can be stopped. When the correct probability value is less than or equal to the preset probability threshold, the relationship semantic detection model is considered not to have completed training, and the cross-entropy loss between the correct probability value and the preset probability threshold can be calculated. The cross-entropy loss function can be used as the loss function of the relationship semantic detection model. For example, the cross-entropy loss function may include:
[0111]
[0112] S328. After optimizing the weights and parameters of the relation semantic detection model according to the cross-entropy loss, retrain the relation semantic detection model.
[0113] In the embodiments of the invention, the weights and parameters of the relation semantic detection model can be optimized according to the cross-entropy loss. After optimizing the weights and parameters, the relation semantic detection model can be retrained according to the above steps until the correct probability value of the predicted label is greater than the preset probability threshold, thus completing the training of the relation semantic detection model.
[0114] In this embodiment of the invention, entity relationships and relationship labels from the training set are input into a relation semantic detection model to obtain semantic vectors corresponding to the training set. A fully connected network is invoked to determine multi-classification for different semantic vectors, generating at least two classification task probability values. The label with the maximum value is selected as the predicted label. When the correct probability value of the predicted label is greater than a preset probability threshold, the entity semantic detection model is considered to have completed training. Otherwise, the cross-entropy loss between the correct probability value and the preset probability threshold is determined. The weights and parameters of the relation semantic detection model are optimized according to the cross-entropy loss, and the relation semantic detection model is retrained. This achieves the training of the relation semantic detection model. By using the cross-entropy loss function as the optimization objective function, the accuracy of the relation semantic detection model is improved, enhancing the user experience.
[0115] Example 4
[0116] Figure 6 This is an architecture diagram for semantic drift detection provided according to Embodiment 4 of the present invention. Figure 6As shown in the diagram, the architecture of semantic drift detection can include a knowledge extraction module, an entity alignment module, and a semantic drift detection module.
[0117] The text data to be identified can include three types: structured data, semi-structured data, and unstructured data. Information can be extracted from the text data through knowledge extraction. Knowledge extraction can include entity extraction and relation extraction.
[0118] In one embodiment, the instruction extraction module may employ a pre-defined knowledge extraction framework for joint extraction, jointly extracting triple information of entities and relations, including the extraction of multiple relations between entities.
[0119] Figure 7 This is a schematic diagram of a preset knowledge extraction framework provided according to Embodiment 4 of the present invention. Figure 7 As shown, the preset knowledge extraction framework may include: a bidirectional long short-term memory network encoder (BiLSTM encoder), an entity recognition module, and a relationship recognition module.
[0120] The BiLSTM Encoder is composed of BiLSTM (Bi-directional Long Short-Term Memory), which is a combination of forward and backward LSTM. The encoded vectors are accumulated using the BiLSTM encoder. In practice, the text data to be recognized is serialized to generate a sequence of text to be recognized, which is then input into the BiLSTM Encoder. Each position in the text data is labeled, and word vectors are determined.
[0121] The entity recognition module is used to automatically discover specific entity text such as device names, organization names, place names, dates, and times. In actual operation, word vectors can be obtained from the BiLSTM Encoder, input into the feedforward neural network, and the softmax loss function is used to calculate the label type probability of each word, thereby extracting specific entity text and entity types.
[0122] The relationship recognition module is used to accumulate the recognized entity vector with the encoded vector processed by the BiLSTM encoder. The encoding result at each position is classified by two classifiers to determine whether it is the start or end position of the entity text. When there are multiple entities in the text to be recognized, the nearest matching principle can be used to pair them. Finally, the entity relationship and the corresponding entity text pair, i.e., the triple, are output.
[0123] In one embodiment, the entity alignment module can be used to determine whether multiple entities in the same or different datasets point to the same entity in the real world, solving the problem of one entity corresponding to multiple names. This solution mainly uses a general entity library (e.g., entity library, thesaurus, etc.) + a domain entity library (e.g., a third-party domain entity library) to complete entity alignment between heterogeneous data through comparison of entity text. In actual operation, the extracted entity text can be retrieved in the index field to obtain candidate entities (candidate entities refer to the text retrieved by Elasticsearch). A score is calculated using the following formula, and a low score threshold is set to filter candidate text: Score = (Number of characters in the intersection of M and Q) / (Number of characters in M). Where M is the candidate entity and Q is the query text. By traversing the query fragments (slicing the query into Q[1:2], Q[1:3], ..., Q[n-1,:n]), the similarity is calculated with each candidate entity using the following formula: Score = 1 - distance(M,P) / (len(M) + len(P)), where M is the candidate entity, P is the query fragment, and distance is the edit distance. The candidate entities are then sorted using the following formula: Score + a * len(P) - b * len(M), where P is the query fragment, M is the candidate entity, a is the matching length weight, and b is the candidate entity length weight. Based on the sorting result, the entity text with the highest score is the aligned entity text.
[0124] In one embodiment, the semantic drift detection module may include a preset semantic drift detection model. The construction of the preset semantic drift detection model may include an entity semantic detection model and a relation semantic detection model. In one embodiment, the preset semantic drift detection model may use a Transformer encoder + Attention + Multi-Tasks as a multi-task learning model to complete two detection tasks: entity semantic detection and relation semantic detection.
[0125] Figure 8 This is a schematic diagram of the structure of a preset semantic drift detection model provided according to Embodiment 4 of the present invention. Figure 8 As shown, the preset semantic drift detection model can include an input layer, a sharing layer, and two task layers.
[0126] The input layer processes the training set. Depending on the task type, there are two input formats: for entity semantic drift detection, the instance input sample format can be entity text, entity text, and entity label, separated by the \t separator. For relation semantic detection, the instance input sample format can be entity relation and relation label, separated by the \t separator.
[0127] In one embodiment, after the training set is input, it can enter the Embedding layer in the sharing layer. After the entity text and entity relations are input into the sharing layer, the positive integer indices of the entity text in the source data can be determined first. The positive integer indices are then transformed using one-hot vectors to capture the relationships between texts and determine the initial text vector and initial relation vector. Then, the encoder in the Transformer determines the entity vector and relation vector corresponding to each entity text in the training set by adding them to the residuals.
[0128] The encoder in the Transformer can be considered as consisting of multiple blocks. Figure 9 This is a schematic diagram of the Transformer encoder provided according to Embodiment 4 of the present invention. Figure 9 As shown, each block adds residual connections, Layer Norm, and Full-Connection (FC) on top of Self-attention, which more accurately determines the entity vectors corresponding to the training set.
[0129] In one embodiment, Figure 10 This is a schematic diagram of the structure of a Block according to Embodiment 4 of the present invention. Specific implementation steps within a single Block may include:
[0130] Step 1: Add the residuals of the original input vector b and the output vector a to obtain the vector a+b;
[0131] Step 2: Obtain vector c from vector a+b using Layer Norm;
[0132] Step 3: Pass vector c through an FC layer to obtain vector d;
[0133] Step 4: Add the residuals of vector c and vector d to obtain vector e;
[0134] Step 5: Vector e is output as vector f through Layer Norm. The output vector f obtained at this time is the output vector of a single block in the Encoder.
[0135] A residual block (shortcut connections / skip connections) is divided into a direct mapping part (xl) and a residual part F(xl, Wl), which can be expressed as: X1 = X1 + (X1, W1). In one embodiment, the Layer Norm calculation formula may include:
[0136] Where E[x] is the expectation and Var[x] is the variance.
[0137] In one embodiment, entity vectors can be fed into an entity semantic detection task. This task employs a Siamese network architecture, taking two entity vectors as input and outputting a real number in the range [0,1], indicating the probability that two entity texts share the same relation type. In actual operation, the preset function may include a distance function, which may include s(x,y)=σ(w s T (f s (x)-f s (y)) 2 +b s In this model, fs(x) and fs(y) represent the output functions of the encoder, where x and y are entity vectors. σ() represents the sigmoid function, ws represents the weights, and bs represents the biases. The mean absolute error loss (MAE) can be used as the objective function for optimization. The entity vectors of two entity texts are input into the preset function, and the similarity value between the two entity vectors is calculated. If the similarity value is greater than a preset similarity threshold, the entity semantic detection model is considered to have completed training. Otherwise, the MAE between the similarity value and the preset similarity threshold is determined, and the entity semantic detection model is retrained after optimizing the weights and parameters according to the MAE.
[0138] In one embodiment, semantic vectors can be input into a relational semantic detection task. This task consists of a linear layer and organizes all relations into a multi-classification problem, where different relations can serve as negative examples of each other. In practice, the semantic vectors can be input into a fully connected network. The fully connected network performs multi-classification on the semantic vectors, determines the probability value for each classification task, and determines the predicted label. In actual operation, when the semantic vectors are input into the fully connected network, the network classifies the semantic vectors and evaluates the predicted value for each classification task, selecting the label with the highest value as the predicted label. Cross-entropy loss can be used as the objective function for optimization. When the correct probability value of the predicted label is greater than a preset probability threshold, the entity semantic detection model is considered to have completed training. Otherwise, the cross-entropy loss between the correct probability value and the preset probability threshold is determined, and the weights and parameters of the relational semantic detection model are optimized according to the cross-entropy loss before retraining the relational semantic detection model.
[0139] In one embodiment, the comprehensive loss function of the preset semantic drift detection model is calculated uniformly from the loss functions of the two tasks mentioned above. The comprehensive loss function can be an operational function used to measure the degree of difference between the predicted value and the true value of the preset semantic drift detection model; the smaller the loss function, the better the robustness of the model. The comprehensive loss function of the preset semantic drift detection model can be determined by the entity semantic detection model and the relation semantic detection model. In one embodiment, the comprehensive loss function can be determined by the entity semantic detection model and the relation semantic detection model. For example, the comprehensive loss function may include:
[0140]
[0141] Where σ1 and σ2 are noise parameters, which control the relative weights of the L1(W) and L2(W) losses respectively. If the noise parameter σ is larger, the weight of the corresponding loss function L(W) will be smaller. However, since the model will try to make the loss function as small as possible, σ will become very large, completely ignoring the influence of the data. Therefore, a regularization term logσ is added to the noise term.
[0142] Example 5
[0143] Figure 11 This is a schematic diagram of a semantic drift detection device provided in Embodiment 5 of the present invention. Figure 11 As shown, the device includes: a text data acquisition module 51, an entity acquisition module 52, and a semantic drift detection module 53.
[0144] The text data acquisition module 51 is used to acquire the text data to be recognized.
[0145] The entity acquisition module 52 is used to acquire the entity type and entity relationship of the entity text in the text data to be identified according to a preset knowledge extraction framework. The knowledge extraction framework includes an entity extraction framework and an entity relationship extraction framework.
[0146] The semantic drift detection module 53 is used to perform semantic drift detection on entity types and entity relationships based on a preset semantic drift detection model to determine the semantic drift situation. The preset semantic drift detection model is trained and generated based on labeled power datasets, power seed sets and unlabeled power data.
[0147] In this embodiment of the invention, a text data acquisition module acquires text data to be identified, an entity acquisition module acquires entity text and entity relationships in the text data to be identified according to a preset knowledge extraction framework, and a semantic drift detection module performs semantic drift detection on entity semantics and entity relationships based on a preset semantic drift detection model to determine the semantic drift situation. This enables convenient detection of semantic drift in power field data, reduces the cost of manual detection, and can eliminate low-quality data to construct a high-quality knowledge graph in the power field.
[0148] In one embodiment, the entity acquisition module 52 includes:
[0149] The part-of-speech tagging unit is used to serialize the text data to be identified into a text sequence to be identified, and then call the bidirectional long short-term memory network in the preset knowledge extraction framework to perform part-of-speech tagging on the text sequence to be identified.
[0150] The type determination unit is used to determine the entity text of the text data to be identified and the entity type of the corresponding entity text according to the softmax loss function of the feedforward neural network in the preset knowledge extraction framework.
[0151] The semantic determination unit is used to call the predictive classifier based on a feedforward neural network within the preset knowledge extraction framework to classify part-of-speech tags and determine the relational semantics corresponding to the entity text in the text data to be identified.
[0152] In one embodiment, a semantic drift detection device further includes:
[0153] The entity alignment module is used to align entity text.
[0154] In one embodiment, the semantic drift detection module 53 includes at least one of the following semantic drift conditions:
[0155] When the similarity value between the entity type in the text data to be identified and the preset entity type is greater than the preset similarity threshold, it is confirmed that the entity type has not undergone semantic drift.
[0156] When the entity relationship output by the preset semantic drift detection model is contained in the text data to be identified, it is confirmed that the entity relationship has not undergone semantic drift.
[0157] If neither the entity type nor the entity relationship has undergone semantic shift, it is confirmed that the text data to be identified has not undergone semantic shift.
[0158] When both entity type and / or entity relationship undergo semantic shift, it is confirmed that the text data to be identified has undergone semantic shift.
[0159] In one embodiment, the semantic drift detection module 53 includes a preset semantic drift detection model comprising an entity semantic detection model and a relation semantic detection model. The preset semantic drift detection model includes an input layer, a shared layer, and two task layers. Accordingly, the training of the preset semantic drift detection model includes:
[0160] Obtain a pre-stored labeled electricity dataset and use it as the training set;
[0161] The training set is input into a pre-built semantic drift detection model for training. The comprehensive loss function of the pre-built semantic drift detection model is determined by the entity semantic detection model and the relation semantic detection model.
[0162] In one embodiment, training the entity semantic detection model includes:
[0163] Input two entity texts and similar labels from the power dataset into the entity semantic detection model to obtain the entity vectors corresponding to the training set.
[0164] Call a preset function to determine the similarity value between entity vectors of different entity texts;
[0165] When the similarity value is greater than the preset similarity threshold, the entity semantic detection model is determined to be trained successfully; otherwise, the average absolute error loss between the similarity value and the preset similarity threshold is determined.
[0166] After optimizing the weights and parameters of the entity semantic detection model using the mean absolute error loss, the entity semantic detection model is retrained.
[0167] In one embodiment, training the relation semantic detection model includes:
[0168] Input the entity relations and relation labels in the training set into the relation semantic detection model to obtain the semantic vectors corresponding to the training set;
[0169] A fully connected network is invoked to determine multi-classification for different feature vectors, generating at least two classification task probability values, and selecting the label with the maximum value as the predicted label.
[0170] If the correct probability value of the predicted label is greater than the preset probability threshold, the entity semantic detection model is determined to be trained successfully; otherwise, the cross-entropy loss between the correct probability value and the preset probability threshold is determined.
[0171] After optimizing the weights and parameters of the relation semantic detection model using cross-entropy loss, the relation semantic detection model is retrained.
[0172] The semantic drift detection device provided in this embodiment of the invention can execute a semantic drift detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0173] Example 6
[0174] Figure 12 This is a schematic diagram of the structure of an electronic device 10 implementing the semantic drift detection method of this invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0175] like Figure 12 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0176] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0177] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a semantic drift detection method.
[0178] In some embodiments, a semantic drift detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the semantic drift detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform a semantic drift detection method by any other suitable means (e.g., by means of firmware).
[0179] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0180] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0181] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0182] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0183] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0184] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0185] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0186] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A semantic drift detection method, characterized in that, include: Obtain the text data to be recognized; The entity types and entity relationships of the entity text in the text data to be identified are obtained according to a preset knowledge extraction framework, wherein the knowledge extraction framework includes an entity extraction framework and an entity relationship extraction framework. The semantic drift detection model is used to detect semantic drift of the entity type and the entity relationship to determine the semantic drift situation. The preset semantic drift detection model is generated by training based on labeled power dataset, power seed set and unlabeled power data. The preset semantic drift detection model includes an entity semantic detection model and a relation semantic detection model. The preset semantic drift detection model includes an input layer, a shared layer, and two task layers. Accordingly, the training of the preset semantic drift detection model includes: Obtain a pre-stored labeled electricity dataset and use the labeled electricity dataset as a training set; The training set is input into the pre-constructed preset semantic drift detection model for training, wherein the comprehensive loss function of the preset semantic drift detection model is determined by the entity semantic detection model and the relation semantic detection model; The training of the entity semantic detection model includes: Input two entity texts and similar labels from the training set into the entity semantic detection model to obtain the entity vectors corresponding to the training set. Call a preset function to determine the similarity value between the entity vectors of different entity texts; When the similarity value is greater than the similarity threshold, the entity semantic detection model is determined to be trained successfully; otherwise, the average absolute error loss between the similarity value and the threshold is determined. After optimizing the weights and parameters of the entity semantic detection model according to the mean absolute error loss, the entity semantic detection model is retrained. The training of the relation semantic detection model includes: Input the entity relations and relation labels in the training set into the relation semantic detection model to obtain the semantic vector corresponding to the training set; A fully connected network is invoked to determine multi-classification for different feature vectors, generating at least two classification task probability values, and selecting the label with the maximum value as the predicted label. When the correct probability value of the predicted label is greater than the probability threshold, the entity semantic detection model is determined to be trained successfully; otherwise, the cross-entropy loss between the correct probability value and the threshold is determined. After optimizing the weights and parameters of the relation semantic detection model using the cross-entropy loss, the relation semantic detection model is retrained.
2. The method according to claim 1, characterized in that, The step of obtaining the entity types and entity relationships of the entity text in the text data to be identified according to the preset knowledge extraction framework includes: The text data to be identified is serialized to generate a text sequence to be identified, and the bidirectional long short-term memory network in the preset knowledge extraction framework is called to perform part-of-speech tagging on the text sequence to be identified. The entity text of the text to be identified and the entity type of the corresponding entity text are determined according to the softmax loss function of the feedforward neural network in the preset knowledge extraction framework. The part-of-speech tagging is classified by a feedforward neural network-based predictive classifier within the preset knowledge extraction framework to determine the relational semantics corresponding to the entity text in the text data to be identified.
3. The method according to claim 1, characterized in that, After obtaining the entity types and entity relationships of the entity text in the text data to be identified according to the preset knowledge extraction framework, the method further includes: Perform entity alignment on the entity text.
4. The method according to claim 1, characterized in that, The semantic drift includes at least one of the following: When the similarity value between the entity type in the text data to be identified and the preset entity type is greater than the preset similarity threshold, it is confirmed that the entity type has not undergone semantic drift. When the entity relationship output by the preset semantic drift detection model is contained in the text data to be identified, it is confirmed that the entity relationship has not undergone semantic drift. If neither the entity type nor the entity relationship has undergone semantic shift, it is confirmed that the text data to be identified has not undergone semantic shift. When semantic shift occurs in both the entity type and / or entity relationship, it is confirmed that semantic shift has occurred in the text data to be identified.
5. A semantic drift detection device, characterized in that, include: The text data acquisition module is used to acquire the text data to be recognized; The entity acquisition module is used to acquire the entity type and entity relationship of the entity text in the text data to be identified according to a preset knowledge extraction framework, wherein the knowledge extraction framework includes an entity extraction framework and an entity relationship extraction framework. The semantic drift detection module is used to perform semantic drift detection on the entity type and the entity relationship based on a preset semantic drift detection model, and determine the semantic drift situation. The preset semantic drift detection model is trained and generated based on labeled power dataset, power seed set and unlabeled power data. The preset semantic drift detection model includes an entity semantic detection model and a relation semantic detection model. The preset semantic drift detection model includes an input layer, a shared layer, and two task layers. Correspondingly, the semantic drift detection device is used for: Obtain a pre-stored labeled electricity dataset and use the labeled electricity dataset as a training set; The training set is input into the pre-constructed preset semantic drift detection model for training, wherein the comprehensive loss function of the preset semantic drift detection model is determined by the entity semantic detection model and the relation semantic detection model; The semantic drift detection device is used for: Input two entity texts and similar labels from the training set into the entity semantic detection model to obtain the entity vectors corresponding to the training set. Call a preset function to determine the similarity value between the entity vectors of different entity texts; When the similarity value is greater than the similarity threshold, the entity semantic detection model is determined to be trained successfully; otherwise, the average absolute error loss between the similarity value and the threshold is determined. After optimizing the weights and parameters of the entity semantic detection model according to the mean absolute error loss, the entity semantic detection model is retrained. The semantic drift detection device is used for: Input the entity relations and relation labels in the training set into the relation semantic detection model to obtain the semantic vector corresponding to the training set; A fully connected network is invoked to determine multi-classification for different feature vectors, generating at least two classification task probability values, and selecting the label with the maximum value as the predicted label. When the correct probability value of the predicted label is greater than the probability threshold, the entity semantic detection model is determined to be trained successfully; otherwise, the cross-entropy loss between the correct probability value and the threshold is determined. After optimizing the weights and parameters of the relation semantic detection model using the cross-entropy loss, the relation semantic detection model is retrained.
6. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the semantic drift detection method according to any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the semantic drift detection method according to any one of claims 1-4.
Citation Information
Patent Citations
Entity relation extraction method and device
CN107784125A
Intelligent learning guide method for Chinese international education
CN109062939A