A semi-supervised device attribute extraction method and system
Through semi-supervised learning and the BERT-Pair-Networks architecture, combined with domain knowledge and confidence-aware pseudo-label optimization, the problems of strong labeling dependence and insufficient domain adaptability in device attribute extraction are solved, and efficient and accurate device attribute extraction is achieved to meet the needs of different device fields.
Patent Information
- Application Number
- CN202510926078.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-07
AI Technical Summary
Existing technologies rely on large amounts of annotated data in device attribute extraction, resulting in high annotation costs and insufficient domain adaptability. It is difficult to effectively handle the professionalism and non-standardization of device description texts. In addition, existing methods do not fully model the semantic relationships between attributes, affecting the structured quality of cataloged data.
By combining a semi-supervised learning method with the BERT-Pair-Networks architecture, a highly robust device attribute extraction model is constructed through domain knowledge-guided initial annotation generation, deep semantic modeling, confidence-aware pseudo-label optimization, and iterative training enhancement. This reduces dependence on manual annotation and improves model generalization capabilities.
It achieves high-precision and high-robust device attribute extraction at low annotation cost, significantly improves the semantic understanding ability of complex text and the flexibility of the model, adapts to different device fields, and supports robust parsing of synonymous expressions and non-standard text.
Smart Images

Figure CN120409483B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of device attribute extraction, and in particular to a semi-supervised device attribute extraction method and system. Background Art
[0002] Device attribute extraction is the process of automatically identifying and extracting key device attribute information (such as model, specifications, performance parameters, and functional characteristics) from unstructured or semi-structured text (such as technical documents, equipment manuals, and research reports). This process is a key step in device cataloging, aiming to provide accurate and comprehensive attribute information for device management and use.
[0003] Traditional methods are primarily based on rules or pattern matching, such as obtaining attribute values through regular expressions and then extracting attribute terms based on dependency parsing. These methods rely on manually formulated rules and templates, and can achieve good results in specific domains and datasets. However, device description texts are often highly specialized, feature diverse terminology, and are often non-standardized (e.g., "A certain brand of electric vehicle has a battery life of >500km"). This makes traditional rule-based methods difficult to generalize, resulting in poor extraction results for complex text and an inability to meet diverse device attribute extraction needs.
[0004] In recent years, with the development of deep learning technology, neural network-based attribute extraction methods have gradually emerged. These methods leverage the semantic perception and generalization learning capabilities of neural networks to automatically learn semantic features in text, thereby improving the accuracy and efficiency of attribute extraction. However, most existing neural network methods are based on supervised learning and require large amounts of labeled data for training. In the device field, labeled data is often limited due to its sensitivity and difficulty in obtaining, which restricts the application of supervised learning methods.
[0005] Semi-supervised learning, a method that combines supervised and unsupervised learning, can be trained on a small amount of labeled data and a large amount of unlabeled data. While reducing the reliance on labeled data, semi-supervised learning methods improve the model's generalization and robustness. However, existing semi-supervised learning methods are rarely used in device attribute extraction, and their advantages have not yet been fully utilized.
[0006] Furthermore, pre-trained language models (such as BERT) have significantly improved text semantic understanding capabilities. BERT-Pair-Networks, an emerging neural network architecture, has demonstrated promising performance in natural language processing tasks such as sentiment analysis. By constructing a joint representation of text pairs, this network can better capture the semantic relationships between texts, providing new insights for device attribute extraction.
[0007] In summary, existing technologies have the following limitations: strong labeling dependence, supervised learning methods require a large amount of domain-labeled data, while equipment domain labeling is costly and time-consuming, and domain adaptability is insufficient. Traditional rule-based methods and general models are not effective in handling equipment professional terms, abbreviations, and polysemous expressions (such as "NX350" may refer to a laptop model or engine parameters); attribute associations are missing, and existing methods do not fully model the semantic relationships between attributes (such as the co-occurrence of "laptop model" and "processor model"), which affects the structured quality of cataloged data.
[0008] Therefore, there is an urgent need for a semi-supervised attribute extraction method for equipment catalog construction. By integrating pre-trained models with domain adaptation technology, high-precision and high-robustness equipment attribute extraction can be achieved at low annotation costs, providing technical support for standardized management of equipment resources. Summary of the Invention
[0009] The purpose of this application is to provide a semi-supervised device attribute extraction method and system, which can achieve high-precision and high-robustness device attribute extraction at low annotation cost.
[0010] To achieve the above objectives, this application provides the following solutions:
[0011] In a first aspect, the present application provides a semi-supervised device attribute extraction method, comprising:
[0012] Get device text;
[0013] identifying key entity information from the device text;
[0014] Generate attribute extraction matching rules based on the key entity information;
[0015] extracting some device attributes based on the attribute extraction matching rule;
[0016] The original text corresponding to some device attributes and the matching template text are used as text pairs to form a seed annotation dataset;
[0017] Build the BERT-Pair-Networks model;
[0018] Determining whether the text pair matches based on the BERT-Pair-Networks model;
[0019] Training the BERT-Pair-Networks model based on the seed annotation dataset;
[0020] Input unlabeled device text data into the pre-trained BERT-Pair-Networks model to generate preliminary prediction results;
[0021] Calculating the confidence level of the preliminary prediction result;
[0022] Select the preliminary prediction results with confidence greater than the preset threshold as pseudo-label samples;
[0023] Integrating the pseudo-label samples into the seed annotation dataset to form new training data;
[0024] Retraining the BERT-Pair-Networks model based on the new training data until the BERT-Pair-Networks model converges to a preset range;
[0025] Extract device attributes based on the trained BERT-Pair-Networks model.
[0026] Optionally, key entity information is identified from the device text, specifically using the following formula:
[0027] ;
[0028] in, represents the set of recognized entities and, represents named entity recognition, Represents device text, are model parameters.
[0029] Optionally, the following formula is specifically used to extract some device attributes based on the attribute extraction matching rule:
[0030] ;
[0031] in, Extract matching rules for device attributes, To extract some device properties, represents the set of recognized entities and, The extracted device attributes are replaced by entities based on matching rules.
[0032] Optionally, the seed annotation dataset is expressed as follows:
[0033] ;
[0034] in, Label the dataset for the seed, The original text extracted for attribute data, The text corresponding to the domain matching rule template, is a label indicating whether the text pair matches, .
[0035] Optionally, determining whether the text pair matches based on the BERT-Pair-Networks model specifically adopts the following formula:
[0036] ;
[0037] in, P Represents the BERT-Pair-Networks model for effective supervised training, are the BERT-Pair-Networks model parameters, The original text extracted for attribute data, The text corresponding to the domain matching rule template, is a label indicating whether the text pair matches, 1 indicates a match. A value of 0 indicates no match.
[0038] Optionally, the BERT-Pair-Networks model is trained based on the seed annotation dataset using a cross entropy loss function:
[0039] ;
[0040] in, Represents the BERT-Pair-Networks model training loss value, represents the seed annotation dataset, Represents text pairs Whether the real label matches, Represents text pairs Whether the predicted label matches, The original text extracted for attribute data, The text corresponding to the field matching rule template.
[0041] Optionally, unlabeled device text data is fed into a pre-trained BERT-Pair-Networks model to generate preliminary predictions using the following formula:
[0042]
[0043] in, represents the preliminary prediction result of the unlabeled text, Indicates whether the BERT-Pair-Networks model predicts whether the text pair matches. Represents unlabeled device text data, The text corresponding to the field matching rule template.
[0044] Optionally, the confidence level of the preliminary prediction result is calculated using the following formula:
[0045] ;
[0046] in, Indicates the confidence of the predicted unlabeled text domain matching rule text, represents the semantic similarity of a text pair, Represents unlabeled device text data, The text corresponding to the field matching rule template.
[0047] Optionally, the convergence condition for the BERT-Pair-Networks model to converge to a preset range is:
[0048] ;
[0049] in, Indicates the convergence judgment of model training, and It represents the F1 value of the evaluation model performance obtained after two complete trainings. Indicates setting a threshold.
[0050] In a second aspect, the present application provides a device attribute extraction system, the device attribute extraction system comprising:
[0051] Device text acquisition module, used to obtain device text;
[0052] A key entity information identification module, configured to identify key entity information from the device text;
[0053] An attribute extraction and matching rule generation module, configured to generate attribute extraction and matching rules based on the key entity information;
[0054] A partial device attribute extraction module, configured to extract partial device attributes based on the attribute extraction matching rule;
[0055] A seed annotation dataset determination module is used to take the original text corresponding to some device attributes and the matching template text as text pairs and form a seed annotation dataset;
[0056] BERT-Pair-Networks model building module, used to build the BERT-Pair-Networks model;
[0057] A text pair matching determination module, configured to determine whether the text pair matches based on the BERT-Pair-Networks model;
[0058] A training module, configured to train the BERT-Pair-Networks model based on the seed annotation dataset;
[0059] The preliminary prediction result generation module is used to input unlabeled device text data into the pre-trained BERT-Pair-Networks model to generate preliminary prediction results;
[0060] A confidence calculation module, used to calculate the confidence of the preliminary prediction result;
[0061] The pseudo-label sample selection module is used to select preliminary prediction results with a confidence level greater than a preset threshold as pseudo-label samples;
[0062] A data fusion module is used to integrate the pseudo-label samples into the seed annotation data set to form new training data;
[0063] A training enhancement module, configured to retrain the BERT-Pair-Networks model based on the new training data until the BERT-Pair-Networks model converges to a preset range;
[0064] The device attribute extraction module is used to extract device attributes based on the trained BERT-Pair-Networks model.
[0065] According to the specific embodiments provided in this application, this application has the following technical effects:
[0066] This application provides a semi-supervised device attribute extraction method and system, which realizes efficient and accurate extraction of device attributes by integrating semi-supervised learning and BERT-Pair-Networks architecture; through domain rule-guided initial annotation generation and deep semantic modeling, it effectively solves the problem of poor generalization ability of traditional rule-based methods caused by strong professionalism and non-standard expressions in the field of device text; adopts domain knowledge-guided initial annotation generation technology to greatly reduce the dependence on manual annotation; based on entity recognition (NER) and rule template matching, it automatically constructs high-quality seed data sets, which significantly improves annotation efficiency and data quality compared with pure manual annotation; based on BERT-Pair-Networks deep The semantic matching model significantly enhances the ability to understand the semantics of complex texts; it learns the fine-grained semantic associations between device text and attributes through the BERT-Pair-Networks network architecture, and can support robust parsing of non-standard texts such as synonymous expressions and omitted sentences; it introduces a confidence perception mechanism to evaluate the reliability of prediction results of unlabeled data; through semantic similarity calculation and confidence threshold screening, it selects high-confidence pseudo-label samples to expand the training set, further optimizes model performance, and improves the robustness and generalization ability of the model; it adopts an iterative optimization training strategy to ensure continuous evolution of the model; through dynamic threshold adjustment, it avoids overfitting while ultimately obtaining a high-performance device attribute extraction model. This strategy can effectively cope with the challenges of model training under small sample labeling conditions. This application can support flexible replacement of domain matching rule libraries and can quickly adapt to different equipment fields such as aviation and weapons. This design not only improves the flexibility and scalability of the system, but also facilitates customized development for different application scenarios. It has broad application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0068] Figure 1 A flowchart of a semi-supervised device attribute extraction method provided in one embodiment of the present application;
[0069] Figure 2 A flowchart of a semi-supervised device attribute extraction method provided in one embodiment of the present application;
[0070] Figure 3 This is a diagram of the BERT-Pair-Networks network architecture in one embodiment of the present application;
[0071] Figure 4 A schematic diagram of the functional modules of a device attribute extraction system provided in one embodiment of the present application;
[0072] Figure 5 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0073] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0074] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0075] Device attribute extraction is an important part of device catalog construction, which aims to provide accurate and comprehensive attribute information for the management and use of devices. However, device description texts have the characteristics of strong domain expertise, diverse terminology, and non-standard expressions, which makes it difficult to generalize traditional rule-based methods, especially when processing complex texts. In addition, most of the existing neural network-based attribute extraction methods rely on a large amount of labeled data for supervised learning, which is difficult to achieve in the field of equipment due to the sensitivity and difficulty of obtaining data. Therefore, this application proposes a semi-supervised device attribute extraction method, specifically a device attribute intelligent extraction method based on semi-supervised learning and BERT-Pair-Networks collaborative optimization. Through the collaborative optimization of four stages: initial annotation generation guided by domain knowledge, BERT-Pair-Networks deep semantic modeling, confidence-aware pseudo-label optimization, and iterative training enhancement, while ensuring model accuracy, it significantly reduces the dependence on manual annotation, and realizes efficient device attribute extraction under small sample annotation conditions.
[0076] First, based on a domain entity library and semantic matching rules, entity recognition (NER) and rule-based template matching techniques are used to extract structured attributes from device texts, constructing a high-quality seed annotated dataset and generating initial annotated data. Secondly, a BERT-Pair-Networks-based attribute extraction model is constructed. The device text and attribute extraction are paired, and semantic associations are learned using the BERT-Pair-Networks architecture. The model is then supervised pre-trained using the initial annotated data, completing the initial training of the BERT-Pair-Networks deep semantic model. The trained model is then applied to a large amount of unlabeled data to generate preliminary predictions. The reliability of the predicted labels is assessed by combining semantic similarity calculation and confidence threshold screening. High-confidence samples are selected as pseudo-annotated data to expand the training set. Finally, the pseudo-annotated data is mixed with the initial annotated data, and the BERT-Pair-Networks model is retrained. Through multiple iterations of optimization, the model performance converges to within a preset threshold (e.g., F1-score fluctuation < threshold), ultimately achieving a highly robust device attribute extraction model.
[0077] This application mainly consists of the following modules: domain knowledge-guided initial annotated data generation module, BERT-Pair-Networks deep semantic attribute extraction module, confidence-aware pseudo-label optimization module, and iterative training enhancement module. The following is an introduction to each of these four modules:
[0078] See also Figure 1 and Figure 2 The method of the present application includes the following steps 101 to 114. Among them:
[0079] The domain knowledge-guided initial annotation data generation module mainly covers steps 101 to 105, which are described in detail as follows:
[0080] Step 101: Get device context.
[0081] Step 102: Identify key entity information from the device context.
[0082] Utilize existing entity recognition (NER) technologies, such as NLTK, HanLP, and SpaCy, combined with device domain knowledge, to identify key entity information from the text, such as place names, dates, times, quantities, percentages, etc. The NER entity recognition results can be expressed as:
[0083] ;
[0084] in, represents the set of recognized entities and, represents named entity recognition, Represents device text, are model parameters.
[0085] Step 103: Generate attribute extraction matching rules based on the key entity information.
[0086] Step 104: extracting some device attributes based on the attribute extraction matching rule.
[0087] Combine the identified entities with the device domain knowledge to generate attribute extraction matching rules. Use regular expressions and other rule-based techniques to extract some attributes. The device attribute extraction matching rules can be expressed as:
[0088] ;
[0089] in, Extract matching rules for device attributes, To extract some device properties, represents the set of recognized entities and, The extracted device attributes are replaced by entities based on matching rules.
[0090] Step 105: The original text corresponding to some device attributes and the matching template text are taken as text pairs to form a seed annotation dataset.
[0091] The original text corresponding to the extracted attribute data and the matching template text are used as text pairs to form a seed annotation dataset. The annotation dataset can be expressed as:
[0092] ;
[0093] in, The original text extracted for attribute data, The text corresponding to the domain matching rule template, is a label indicating whether the text pair matches, .
[0094] The BERT-Pair-Networks deep semantic attribute extraction module mainly covers steps 106 to 108, which are described as follows:
[0095] Step 106: Build the BERT-Pair-Networks model.
[0096] Step 107: Determine whether the text pair matches based on the BERT-Pair-Networks model.
[0097] Build an attribute extraction model based on BERT-Pair-Networks. For the specific structure, see Figure 3, the model predicts whether a text pair matches by learning the semantic association between the text pairs. The model can be expressed as:
[0098] ;
[0099] in, are the BERT-Pair-Networks model parameters, The original text extracted for attribute data, The text corresponding to the domain matching rule template, is a label indicating whether the text pair matches, 1 indicates a match. A value of 0 indicates no match.
[0100] Step 108: Train the BERT-Pair-Networks model based on the seed annotation dataset.
[0101] Using seed annotation dataset Perform supervised pre-training on the BERT-Pair-Networks model to optimize the model parameters. The loss function can be expressed as:
[0102] ;
[0103] in, Represents the BERT-Pair-Networks model training loss value, represents the seed annotation dataset, Represents text pairs Whether the real label matches, Represents text pairs Whether the predicted label matches, The original text extracted for attribute data, The text corresponding to the field matching rule template.
[0104] Precision, recall, and F1-score are used as performance evaluation indicators:
[0105] Precision refers to the proportion of samples that are actually positive among all samples predicted to be positive. The calculation formula is:
[0106] ;
[0107] Recall refers to the proportion of samples that are correctly predicted to be positive among all samples that are actually positive. The calculation formula is:
[0108] ;
[0109] The F1 score is the harmonic mean of precision and recall, which takes into account the balance between precision and recall. The calculation formula is:
[0110] ;
[0111] Among them, TP (True Positive) represents the number of samples correctly predicted as positive, FP (False Positive) represents the number of samples incorrectly predicted as positive, and FN (False Negative) represents the number of samples incorrectly predicted as negative.
[0112] The performance evaluation results can be used to fine-tune the parameters of the BERT-Pair-Networks model. For example, the learning rate and the weight of the loss function can be adjusted.
[0113] The confidence-aware pseudo-label optimization module mainly covers steps 109 to 112, which are described in detail as follows:
[0114] Step 109: Input the unlabeled device text data into the pre-trained BERT-Pair-Networks model to generate preliminary prediction results.
[0115] Large amounts of unlabeled device text data Input the pre-trained BERT-Pair-Networks model to generate preliminary prediction results :
[0116]
[0117] in, represents the preliminary prediction result of the unlabeled text, Indicates whether the BERT-Pair-Networks model predicts whether the text pair matches. Represents unlabeled device text data, The text corresponding to the field matching rule template.
[0118] Step 110: Calculate the confidence level of the preliminary prediction result.
[0119] Combining semantic similarity calculation and confidence threshold screening, the reliability of the predicted label is evaluated. The confidence can be expressed as:
[0120] ;
[0121] in, Indicates the confidence of the predicted unlabeled text domain matching rule text, represents the semantic similarity of a text pair, Represents unlabeled device text data, The text corresponding to the field matching rule template.
[0122] Step 111: Select the preliminary prediction results with a confidence level greater than a preset threshold as pseudo-label samples.
[0123] That is, only the confidence scores above the threshold are retained The pseudo-label sample is as follows:
[0124]
[0125] in, Represents high-confidence pseudo-label samples.
[0126] Step 112: Integrate the pseudo-label samples into the seed annotation data to form new training data.
[0127] That is, the filtered high confidence pseudo-label samples Add training set and expand the scale of labeled data. The new data set can be expressed as:
[0128] .
[0129] The iterative training enhancement module mainly covers steps 113 and 114, which are described in detail as follows:
[0130] Step 113: Retrain the BERT-Pair-Networks model based on the new training data until the BERT-Pair-Networks model converges to a preset range.
[0131] Step 114: Extract device attributes based on the trained BERT-Pair-Networks model.
[0132] Through multiple iterations of training, until the model performance converges to the preset threshold range (such as F1-score fluctuation is less than the set threshold ), and finally obtain a high-performance device attribute extraction model. The convergence condition can be expressed as:
[0133] ;
[0134] in, Indicates the convergence judgment of model training, and It represents the F1 value of the evaluation model performance obtained after two complete trainings. Indicates setting a threshold.
[0135] In order to facilitate a further understanding of the present application, the present application is further described in detail below using a specific example. The specific implementation process is as follows:
[0136] Text 1:
[0137] Vehicle name: Volkswagen Tiguan L, Manufacturer: SAIC Volkswagen, Production date: December 2022, Storage location: First parking lot, Vehicle model: Tiguan L 2023, Engine type: 1.5T, Fuel consumption: 6.5L / 100km, Vehicle color: Gray, Vehicle configuration: Smart version.
[0138] Text 2:
[0139] The Volkswagen Magotan in the second parking lot was meticulously built by FAW-Volkswagen in March 2023. Equipped with a 1.4T engine, it boasts excellent fuel efficiency, consuming only 5.9 liters per 100 kilometers. The vehicle is primarily white, and its features are designed for comfort.
[0140] Text 3:
[0141] In the new energy vehicle parking lot, there is a car with a very advanced power system and excellent mileage. The body color is very stylish and the configuration is also very high-end, which is very suitable for environmentally friendly travel.
[0142] 1. Domain Knowledge-Guided Initial Annotated Data Generation
[0143] 1.1 Entity Recognition
[0144] Using SpaCy and government vehicle domain knowledge, we can identify key entity information from text.
[0145] For example, in text 1, T = "Volkswagen Tiguan L, Manufacturer: SAIC Volkswagen, Production Date: December 2022, Storage Location: Parking Lot 1, Vehicle Model: Tiguan L 2023, Engine Type: 1.5T, Fuel Consumption: 6.5L / 100km, Vehicle Color: Gray, Vehicle Configuration: Smart Edition."
[0146] The result after entity recognition is:
[0147] E=[('Volkswagen Tiguan L', 'PRODUCT'), ('SAIC Volkswagen', 'ORG'), ('December 2022', 'DATE'), ('First Parking Lot', 'LOC'), ('Tiguan L 2023 Model', 'PRODUCT'), ('1.5T', 'PRODUCT'), ('6.5L / 100km', 'QUANTITY')]
[0148] 1.2 Building domain matching rules
[0149] Generate attribute extraction matching rules based on the identified entities and device domain knowledge.
[0150] For example, define the rule:
[0151] R = {Name [PRODUCT] Manufacturer [ORG] Production Date [DATE] Storage Location [LOC] Engine Type [PRODUCT] Fuel Consumption [QUANTITY]}
[0152] Extract structured attributes using regular expressions:
[0153] A=[('Name', Volkswagen Tiguan L), ('Generation Date', December 2022), ('Location', Parking Lot 1), ('Engine', '1.5T'), ('Fuel Consumption', '6.5L / 100km')]
[0154] 1.3 Construction of labeled dataset
[0155] D=[("Name [PRODUCT] Manufacturer [ORG] Production Date [DATE] Storage Location [LOC] Engine Type [PRODUCT] Fuel Consumption [QUANTITY]", "Vehicle Name: [PRODUCT], Manufacturer: [ORG], Production Date: [DATE], Storage Location: [LOC], Vehicle Model: Tiguan L 2023, Engine Type: [PRODUCT], Fuel Consumption: [QUANTITY], Vehicle Color: Gray, Vehicle Configuration: Smart Edition.", 1),……]
[0156] 2. BERT-Pair-Networks Deep Semantic Attribute Extraction
[0157] Use the initial labeled dataset D to perform supervised pre-training on the BERT-Pair-Networks model and optimize the model parameters. The loss function can be expressed as:
[0158] .
[0159] 2.3 Performance Evaluation
[0160] Precision, recall, and F1-score are used as performance evaluation indicators:
[0161] Precision: Precision refers to the proportion of samples that are actually positive among all samples predicted to be positive. The calculation formula is:
[0162] .
[0163] Among them, TP (True Positive) represents the number of samples correctly predicted as positive, and FP (False Positive) represents the number of samples incorrectly predicted as positive.
[0164] Recall: Recall refers to the proportion of samples that are correctly predicted to be positive among all samples that are actually positive. The calculation formula is:
[0165] .
[0166] Among them, FN (False Negative) represents the number of samples that are incorrectly predicted as negative.
[0167] F1-Score: The F1 score is the harmonic mean of precision and recall, which comprehensively considers the balance between precision and recall.
[0168] The calculation formula is:
[0169] .
[0170] 3. Confidence-aware pseudo-label optimization
[0171] 3.1 Unlabeled Data Prediction
[0172] For example, in text 2, T="This Volkswagen Magotan in the second parking lot is a 2023 model meticulously built by FAW-Volkswagen in March 2023. Equipped with a 1.4T engine, it boasts excellent fuel efficiency, consuming only 5.9 liters per 100 kilometers. The vehicle's exterior is primarily white, and its features are designed for comfort."
[0173] The text pairs that need to be predicted are:
[0174] ("Name [PRODUCT] Manufacturer [ORG] Production Date [DATE] Storage Location [LOC] Engine Type [PRODUCT] Fuel Consumption [QUANTITY]", "This [PRODUCT] located at [LOC] was carefully built by FAW-Volkswagen on [DATE]. Equipped with a [PRODUCT] engine, it has excellent fuel efficiency, consuming only [QUANTITY] per 100 kilometers. The vehicle's exterior is mainly white, and its configuration is designed for comfort."), prediction result 1 indicates a match, and the following attributes are extracted: name, manufacturer, production date, storage location, engine type, and fuel consumption.
[0175] For example, in text 3, T="In the new energy vehicle parking lot, there is a car with a very advanced power system and excellent mileage. The car body color is very stylish and the configuration is also very high-end, making it very suitable for environmentally friendly travel."
[0176] The text pairs that need to be predicted are:
[0177] ("Name [PRODUCT] Manufacturer [ORG] Production Date [DATE] Storage Location [LOC] Engine Type [PRODUCT] Fuel Consumption [QUANTITY]", "In [LOC], there is a car with a very advanced power system and excellent mileage. The car body color is very fashionable and the configuration is also very high-end, which is very suitable for environmentally friendly travel"), the prediction result is 0, indicating no match, and no attributes are extracted.
[0178] 3.2 Confidence Assessment
[0179] Combining semantic similarity calculation and confidence threshold screening, the reliability of the predicted label is evaluated. The confidence can be expressed as:
[0180] .
[0181] Only keep the confidence scores above the threshold Pseudo-labeled examples:
[0182] .
[0183] Calculate the semantic similarity of text pairs:
[0184] For example, if the semantic similarity of the text pair ("Name [PRODUCT] Manufacturer [ORG] Production Date [DATE] Storage Location [LOC] Engine Type [PRODUCT] Fuel Consumption [QUANTITY]", "This [PRODUCT] located in [LOC] was carefully built by FAW-Volkswagen on [DATE]. Equipped with a [PRODUCT] engine, it has excellent fuel efficiency, consuming only [QUANTITY] per 100 kilometers. The vehicle's exterior is mainly white, and its configuration is designed for comfort.") is 0.99, and the confidence threshold is 0.98, then the sample ("Name [PRODUCT] Manufacturer [ORG] Production Date [DATE] Storage Location [LOC] Engine Type [PRODUCT] Fuel Consumption [QUANTITY]", "This [PRODUCT] at [LOC] was meticulously built by FAW-Volkswagen on [DATE]. Equipped with a [PRODUCT] engine, it boasts excellent fuel efficiency, consuming only [QUANTITY] per 100 kilometers. The vehicle's exterior is primarily white, and its features are designed for comfort." 1) is a pseudo-label sample with a confidence level above the threshold.
[0185] 3.3 Pseudo-label expansion
[0186] The filtered high-confidence pseudo-label samples are added to the training set to expand the scale of labeled data.
[0187] 4. Iterative Training Enhancement Module
[0188] 4.1 Incorporating Pseudo-label Hybrid Training
[0189] The pseudo-labeled data is mixed with the initial annotated data and reused for training the BERT-Pair-Networks model.
[0190] 4.2 Iterative Optimization
[0191] Through multiple iterations of training, until the model performance converges to the preset threshold range (such as F1-score fluctuation is less than the set threshold ), and finally obtain a high-performance device attribute extraction model. The convergence condition can be expressed as:
[0192] .
[0193] For example, if the F1-score fluctuation between two adjacent iterations is less than the threshold 0.01, the iteration is stopped.
[0194] Based on the same inventive concept, embodiments of the present application also provide a device attribute extraction system for implementing the aforementioned semi-supervised device attribute extraction method. The solution provided by this system is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more device attribute extraction system embodiments provided below can be found in the aforementioned limitations of the semi-supervised device attribute extraction method and will not be further elaborated here.
[0195] In an exemplary embodiment, Figure 4 As shown, a device attribute extraction system is provided, comprising:
[0196] Device text acquisition module 201, used to acquire device text;
[0197] A key entity information identification module 202, configured to identify key entity information from the device text;
[0198] An attribute extraction and matching rule generation module 203 is configured to generate attribute extraction and matching rules based on the key entity information;
[0199] A partial device attribute extraction module 204 is configured to extract partial device attributes based on the attribute extraction matching rule;
[0200] A seed annotation dataset determination module 205 is configured to take the original text corresponding to some device attributes and the matching template text as text pairs and form a seed annotation dataset;
[0201] A BERT-Pair-Networks model construction module 206, for constructing a BERT-Pair-Networks model;
[0202] A text pair matching determination module 207 is configured to determine whether the text pair matches based on the BERT-Pair-Networks model;
[0203] A training module 208 is configured to train the BERT-Pair-Networks model based on the seed annotation dataset;
[0204] A preliminary prediction result generation module 209 is used to input unlabeled device text data into the pre-trained BERT-Pair-Networks model to generate preliminary prediction results;
[0205] A confidence calculation module 210 is used to calculate the confidence of the preliminary prediction result;
[0206] The pseudo-label sample selection module 211 is used to select preliminary prediction results with a confidence level greater than a preset threshold as pseudo-label samples;
[0207] A data fusion module 212 is used to integrate the pseudo-label samples into the seed annotation data set to form new training data;
[0208] A training enhancement module 213 is configured to retrain the BERT-Pair-Networks model based on the new training data until the BERT-Pair-Networks model converges to a preset range;
[0209] The device attribute extraction module 214 is used to extract device attributes based on the trained BERT-Pair-Networks model.
[0210] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store device attribute extraction data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a semi-supervised device attribute extraction method is implemented.
[0211] Those skilled in the art will understand that Figure 5 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0212] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0213] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0214] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0215] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0216] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0217] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.
[0218] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0219] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A semi-supervised device attribute extraction method, characterized in that: The semi-supervised device attribute extraction method comprises: Get device text; identifying key entity information from the device text; Generate attribute extraction matching rules based on the key entity information; extracting some device attributes based on the attribute extraction matching rule; The original text corresponding to some device attributes and the matching template text are used as text pairs to form a seed annotation dataset; Build the BERT-Pair-Networks model; Determining whether the text pair matches based on the BERT-Pair-Networks model; Training the BERT-Pair-Networks model based on the seed annotation dataset; Input unlabeled device text data into the pre-trained BERT-Pair-Networks model to generate preliminary prediction results; Calculating the confidence level of the preliminary prediction result; Select the preliminary prediction results with confidence greater than the preset threshold as pseudo-label samples; Integrating the pseudo-label samples into the seed annotation dataset to form new training data; Retraining the BERT-Pair-Networks model based on the new training data until the BERT-Pair-Networks model converges to a preset range; Extract device attributes based on the trained BERT-Pair-Networks model; The confidence level of the preliminary prediction result is calculated using the following formula: ; in, Indicates the confidence of the predicted unlabeled text domain matching rule text, represents the semantic similarity of a text pair, Represents unlabeled device text data, The text corresponding to the domain matching rule template; The convergence conditions for the BERT-Pair-Networks model to converge to the preset range are: ; in, Indicates the convergence judgment of model training, and It represents the F1 value of the evaluation model performance obtained after two complete trainings. Indicates setting threshold; The following formula is used to extract some device attributes based on the attribute extraction matching rule: ; in, Extract matching rules for device attributes, To extract some device properties, represents the set of recognized entities and, To extract device attributes through entity replacement based on matching rules; The following formula is specifically used to determine whether the text pair matches based on the BERT-Pair-Networks model: ; in, P represents the BERT-Pair-Networks model, are the BERT-Pair-Networks model parameters, The original text extracted for attribute data, The text corresponding to the domain matching rule template, is a label indicating whether the text pair matches, 1 indicates a match. A value of 0 indicates no match.
2. The semi-supervised device attribute extraction method according to claim 1, characterized in that Identify key entity information from the device text using the following formula: ; in, represents the set of recognized entities and, represents named entity recognition, Represents device text, are model parameters.
3. The semi-supervised device attribute extraction method according to claim 1, characterized in that The expression of the seed annotation dataset is as follows: ; in, Label the dataset for the seed, The original text extracted for attribute data, The text corresponding to the domain matching rule template, is a label indicating whether the text pair matches, .
4. The semi-supervised device attribute extraction method according to claim 1, characterized in that The BERT-Pair-Networks model is trained based on the seed annotation dataset using the cross entropy loss function: ; in, Represents the BERT-Pair-Networks model training loss value, represents the seed annotation dataset, Represents text pairs Whether the real label matches, Represents text pairs Whether the predicted label matches, The original text extracted for attribute data, The text corresponding to the field matching rule template.
5. The semi-supervised device attribute extraction method according to claim 1, characterized in that The unlabeled device text data is fed into the pre-trained BERT-Pair-Networks model to generate preliminary prediction results using the following formula: ; in, represents the preliminary prediction result of the unlabeled text, Indicates whether the BERT-Pair-Networks model predicts whether the text pair matches. Represents unlabeled device text data, The text corresponding to the field matching rule template.
6. A device attribute extraction system, characterized in that: The device attribute extraction system includes: Device text acquisition module, used to obtain device text; A key entity information identification module, configured to identify key entity information from the device text; An attribute extraction and matching rule generation module, configured to generate attribute extraction and matching rules based on the key entity information; A partial device attribute extraction module, configured to extract partial device attributes based on the attribute extraction matching rule; A seed annotation dataset determination module is used to take the original text corresponding to some device attributes and the matching template text as text pairs and form a seed annotation dataset; BERT-Pair-Networks model building module, used to build the BERT-Pair-Networks model; A text pair matching determination module, configured to determine whether the text pair matches based on the BERT-Pair-Networks model; A training module, configured to train the BERT-Pair-Networks model based on the seed annotation dataset; The preliminary prediction result generation module is used to input unlabeled device text data into the pre-trained BERT-Pair-Networks model to generate preliminary prediction results; A confidence calculation module, used to calculate the confidence of the preliminary prediction result; The pseudo-label sample selection module is used to select preliminary prediction results with a confidence level greater than a preset threshold as pseudo-label samples; A data fusion module is used to integrate the pseudo-label samples into the seed annotation data set to form new training data; A training enhancement module, configured to retrain the BERT-Pair-Networks model based on the new training data until the BERT-Pair-Networks model converges to a preset range; Device attribute extraction module, used to extract device attributes based on the trained BERT-Pair-Networks model; The confidence level of the preliminary prediction result is calculated using the following formula: ; in, Indicates the confidence of the predicted unlabeled text domain matching rule text, represents the semantic similarity of a text pair, Represents unlabeled device text data, The text corresponding to the domain matching rule template; The convergence conditions for the BERT-Pair-Networks model to converge to the preset range are: ; in, Indicates the convergence judgment of model training, and It represents the F1 value of the evaluation model performance obtained after two complete trainings. Indicates setting threshold; The following formula is used to extract some device attributes based on the attribute extraction matching rule: ; in, Extract matching rules for device attributes, To extract some device properties, represents the set of recognized entities and, To extract device attributes through entity replacement based on matching rules; The following formula is specifically used to determine whether the text pair matches based on the BERT-Pair-Networks model: ; in, P represents the BERT-Pair-Networks model, are the BERT-Pair-Networks model parameters, The original text extracted for attribute data, The text corresponding to the domain matching rule template, is a label indicating whether the text pair matches, 1 indicates a match. A value of 0 indicates no match.
Citation Information
Patent Citations
Entity analysis method and device, medium and product
CN119378550A