Metadata extraction method and device, electronic equipment and storage medium
By obtaining the feature vectors of text block data of document data, using the feature vectors to obtain candidate metadata and further obtain target metadata, the problems of document format diversity and inconsistent professional terminology in the power industry are solved, the accuracy of metadata extraction is improved, and the intelligent decision-making capabilities of power grid operation monitoring and equipment maintenance are enhanced.
Patent Information
- Application Number
- CN202510739232.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-12
AI Technical Summary
The diversity of document formats and inconsistent professional terminology generated by different departments in the power industry lead to low metadata extraction accuracy.
By obtaining the feature vector of the text block data of the document data, the feature vector is used to obtain candidate metadata, and the target metadata is obtained based on the candidate metadata, which can be adapted to metadata extraction from document data with different formats and inconsistent professional terminology.
It improves the accuracy of metadata extraction, adapts to document data in different formats and with inconsistent professional terminology, and enhances the level of intelligent decision-making in power grid operation monitoring, equipment maintenance, and fault warning.
Smart Images

Figure CN120632068A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data management technology, for example, to a method and device, electronic device, and storage medium for metadata extraction. Background Art
[0002] Currently, the power industry is plagued by a vast number of documents generated by various departments, including grid operation reports, equipment maintenance records, fault detection reports, and load forecast reports. These documents exhibit significant differences in format, structure, and terminology, making it difficult to automatically extract and integrate key information. Traditional information extraction methods based on rules and statistical models, while achieving high accuracy for documents with fixed formats, often struggle to adapt to the diverse formats and inconsistent terminology. This results in low metadata accuracy when extracting data from power industry documents.
[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0004] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0005] The embodiments of the present disclosure provide a method and apparatus, an electronic device, and a storage medium for metadata extraction, so as to improve the accuracy of metadata extraction.
[0006] In some embodiments, the method includes: obtaining document data to be extracted; obtaining text block data of the document data to be extracted; obtaining a feature vector of the text block data; obtaining candidate metadata based on the feature vector; and obtaining target metadata of the document data to be extracted based on the candidate metadata.
[0007] In some embodiments, the device includes: a document data acquisition module, configured to acquire document data to be extracted; a text block data acquisition module, configured to acquire text block data of the document data to be extracted; a feature vector acquisition module, configured to acquire feature vectors of the text block data; a candidate metadata acquisition module, configured to acquire candidate metadata based on the feature vector; and a target metadata acquisition module, configured to acquire target metadata of the document data to be extracted based on the candidate metadata.
[0008] In some embodiments, the electronic device includes: a processor and a memory storing program instructions, and the processor is configured to execute the above-mentioned method for metadata extraction when running the program instructions.
[0009] In some embodiments, the storage medium stores program instructions, and the program instructions are executed by a processor to implement the above-mentioned method for metadata extraction.
[0010] The method and apparatus, electronic device, and storage medium for metadata extraction provided by the embodiments of the present disclosure can achieve the following technical effects: The system can obtain feature vectors of text blocks from the document data to be extracted, first obtain candidate metadata based on the feature vectors of the text blocks, and then obtain target metadata based on the candidate metadata. This method can adapt to metadata extraction from documents with different formats and inconsistent professional terminology, and further obtain target metadata based on the candidate metadata, thereby achieving higher accuracy in the obtained target metadata.
[0011] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] One or more embodiments are exemplarily described by corresponding drawings. These exemplary descriptions and drawings do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation. In addition, Figure 1 is a schematic diagram of a method for metadata extraction provided by an embodiment of the present disclosure; Figure 2 is a schematic diagram of a device for metadata extraction provided by an embodiment of the present disclosure; Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0013] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The accompanying drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.
[0014] In the description and claims of the embodiments of the present disclosure, as well as in the accompanying drawings, the terms "first," "second," and the like are used to distinguish similar items and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate to describe the embodiments of the present disclosure herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions.
[0015] Unless otherwise stated, the term "plurality" means two or more.
[0016] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.
[0017] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0018] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.
[0019] The disclosed embodiment provides a method for metadata extraction. The execution subject is an electronic device, which includes a computer or server. The electronic device extracts feature vectors of text block data from the document data to be extracted, obtains candidate metadata based on the feature vectors, and further obtains target metadata based on the candidate metadata. The method is capable of adapting to metadata extraction from document data of different formats and with inconsistent professional terminology, thereby increasing the accuracy of the target metadata obtained. The method realizes the automatic extraction of multidimensional metadata from heterogeneous documents in the power industry, thereby improving the level of intelligent decision-making for power grid operation monitoring, equipment maintenance, and fault warning.
[0020] Combine Figure 1 As shown, the embodiment of the present disclosure provides a method for metadata extraction, including: Step S101: Obtain document data to be extracted. The document data to be extracted is document data from which metadata needs to be extracted. For example, document data to be extracted in the power industry may include power grid operation reports, equipment maintenance records, fault detection reports, and load forecast reports.
[0021] Step S102: obtaining text block data of the document data to be extracted.
[0022] Step S103: obtaining a feature vector of the text block data.
[0023] Step S104: Obtain candidate metadata based on the characteristic vector of the text block data.
[0024] Step S105 : acquiring target metadata of the document data to be extracted based on the candidate metadata.
[0025] The metadata extraction method provided by the embodiments of the present disclosure can obtain feature vectors of text blocks in the document data to be extracted, first obtain candidate metadata based on the feature vectors of the text blocks, and then obtain target metadata based on the candidate metadata. Obtaining candidate metadata based on the feature vectors of the text blocks in the document data to be extracted allows metadata extraction from documents of varying formats and with inconsistent terminology. Furthermore, the target metadata can be obtained based on the candidate metadata, resulting in higher accuracy of the obtained target metadata.
[0026] Furthermore, obtaining text block data of the document data to be extracted includes: extracting text information from the document data to be extracted, dividing the text information into text blocks, and obtaining text block data.
[0027] In some embodiments, extracting text information from document data to be extracted includes: obtaining the file extension and file header information of the document data to be extracted. Determining the document format of the document data to be extracted based on the file extension and file header information. Calling the corresponding parsing module according to the document format to convert the document data to be extracted into text information. For example, for document data to be extracted in formats such as PDF, Word, and XML, calling the corresponding parsing module to convert the document data to be extracted into text information. During the text format conversion process, the corresponding parsing tool can be used to accurately identify various document structures, ensuring that the text content is complete and without omission. At the same time, regular expressions and preset template matching are used to extract structural information such as headers, footers, tables, and titles in each text information.
[0028] In some embodiments, when the document data to be extracted is in the form of a scanned document, text information is extracted using OCR (Optical Character Recognition) technology to ensure that the text in the image can be accurately converted into text information.
[0029] In some embodiments, pre-written regular expressions and template rules are used to remove noise information from text information. For example, noise information such as headers, footers, page numbers, chart captions, and copyright information is removed. In addition, the noise rate of each text information is ensured to be less than a preset value, for example, the preset value is 2.5%. That is, η is less than 2.5%, where η is the noise rate. By calculating Obtain the noise rate, where for , L is the total number of characters in the text information. This can provide high-quality text information for subsequent analysis, thereby further improving the accuracy of metadata extraction.
[0030] Optionally, after extracting text information from the document data to be extracted, the method further includes normalizing the preset terms in the text information, thereby unifying the professional terms in the text information and improving the accuracy of subsequent metadata extraction.
[0031] Specifically, by calculating: Obtain the standard term after normalization. is a preset term in the text message, Normalize() is a function that maps pre-set terms in text information to unified standard terms based on a preset dictionary and synonym library, ensuring consistent terminology across modules during subsequent processing. The preset dictionary contains standard terms. For example, the power industry standard terminology dictionary includes terms like "circuit breaker," "busbar," "transformer," and "equipment number." The synonym library is a synonym mapping library for standard terms.
[0032] Furthermore, dividing the text information into text blocks to obtain text block data includes: obtaining a structural identifier in the text information, and dividing the text information into at least one text block data according to the structural identifier. For example, the structural identifier includes one or more of: a main title, a subtitle, a blank line, and a special identifier.
[0033] In some embodiments, the text information is segmented into text blocks using a preset automatic segmentation algorithm using structured information of the text information to obtain text block data. For example, the structured information includes one or more of a title, body, table, and appendix.
[0034] The number of words in each text block is controlled between 80 and 120 words to ensure that the text block contains sufficient context information, while avoiding excessive length that may dilute the information or make it difficult to process.
[0035] Furthermore, obtaining a feature vector for the text block data includes: obtaining contextual features of the text block data, obtaining a matching score between a preset keyword and the text block data, obtaining structural features of the text block data, and obtaining a frequency of each word in the text block data and a TF-IDF value of each word. The feature vector for the text block data is obtained based on the contextual features of the text block data, the matching score between the preset keyword and the text block data, the structural features of the text block data, the frequency of each word in the text block data, and the TF-IDF value of each word.
[0036] The context features of the extracted text block data, the matching scores between the preset keywords and the text block data, the structural features of the text block data, the frequency of each word in the text block data and the TF-IDF value of each word are integrated to form the feature vector of the text block data. , where d is the total dimension of the feature vector. Obtaining candidate metadata based on the feature vector can improve the accuracy of obtaining candidate metadata.
[0037] Furthermore, the context features of the text block data are obtained, including: extracting the first text information of a set number of lines before the text block data, and extracting the second text information of a set number of lines after the text block data. The first text information, the text block data, and the second text information are merged to obtain a context string. The context string is segmented and stop words or punctuation marks are removed. The word frequency of each word in the context string is obtained, and the word frequency of each word is mapped to a vector space according to a pre-set vocabulary or feature selection method to obtain the context features of the text block data. Each dimension corresponds to a word, and the value is the number of times the word appears in the context string. The context features of the text block data can reflect its context information. For example, the number of lines is set to 3. If it is less than the set number of lines, the actual number of lines is taken.
[0038] Furthermore, obtaining a matching score between a preset keyword and the text block data includes: obtaining an occurrence frequency of the preset keyword in the text block data, and determining the obtained occurrence frequency as a matching score between the corresponding keyword and the text block data.
[0039] The preset keywords are obtained by obtaining the industry domain to which the document data corresponding to the text block data belongs. Using the industry domain, a table lookup operation is performed in a preset first data table to find keywords corresponding to the industry domain. The first data table stores the correspondence between industry domains and keywords. For example, if the industry domain is the power industry, the keywords found include "fault," "shutdown," and "maintenance." Thus, the matching score between the preset keywords and the text block data can reflect the relevance of the text block data to the professional field of the power industry.
[0040] Furthermore, obtaining the structural feature of the text block data includes: obtaining the position of the text block data in the corresponding document data, and determining the position feature as the structural feature of the text block data.
[0041] Furthermore, the frequency of each word in the text block data is obtained, including: using a preset word segmentation tool to segment the text block data, and calculating Get the term frequency (TF) of each word in the text block data, where For the j Text block data Middle i words The word frequency, For the j The first i The number of occurrences of a word, For the j The total number of words in a text block. The default word segmentation tools include jieba or HanLP, etc.
[0042] Further, the TF-IDF value of each word in the text block data is obtained, including: by calculating Get the TF-IDF value of each word in the text block data. For the j The first i The TF-IDF value of the word, For the i words The Inverse Document Frequency (IDF) of
[0043] Among them, by calculating Get the first i The inverse document frequency of the word. is the number of document data, For the word containing i The number of documents containing the word.
[0044] Furthermore, candidate metadata is obtained based on the feature vector, including: inputting the feature vector of the text block data into a preset classification model to obtain the candidate metadata. Specifically, the feature vector of the text block data is input into the preset classification model, and the classification model outputs the candidate metadata and a classification score corresponding to each candidate metadata. The classification score represents the confidence level that the corresponding text block data belongs to the candidate metadata. For example, the candidate metadata output by the classification model may be "equipment number," "fault description," "maintenance date," or "maintenance personnel."
[0045] Optionally, the preset classification model is obtained by: obtaining multiple sets of sample document data, obtaining sample text blocks from each set of sample document data, labeling each set of sample text blocks, where the labels include candidate metadata and classification scores corresponding to each set of sample text blocks. Obtaining a sample feature vector for each set of sample text blocks. Inputting each sample feature vector into a preset support vector machine (SVM) for training to obtain a classification model. In one embodiment, 2,000 sets of sample document data are obtained.
[0046] In some embodiments, a radial basis function (RBF) is used as a kernel function to capture nonlinear relationships. The regularization coefficient C and kernel parameter γ of the SVM are determined through 4-fold cross-validation. The goal of cross-validation is to minimize the validation error, and its objective function is: .in, is the classification error rate during the k-fold cross-validation. After parameter optimization, the SVM classification accuracy is greater than or equal to 92%.
[0047] In some embodiments, each candidate metadata is defined as a state q, such as "device number", "fault description", etc. The training set statistics state q is transferred to state The number of times c, by calculating Get state q and transfer to state The probability of , where Transition from state q to state The probability of Transition from state q to state The number of times, represents the total number of transitions from state q to all possible successor states q''. This is the sum of the number of transitions immediately following the appearance of state q in the corpus (or observation sequence). This value serves as the normalized denominator for calculating the transition probabilities from q to each state q''. Thus, calculating transition probabilities between states reflects the relative frequency and order of occurrence of candidate metadata in the document data.
[0048] In some embodiments, the position of the word in the text block in the corresponding state is different, and the method of obtaining the emission probability is also different. Located at the first position of state q, by calculating Get the emission probability ,in, is the word in state q Number of occurrences. For words The previous word of . Indicates that in the corpus, the word Previous word As a condition, the total number of times state q appears, that is, for all possible second preceding words on The accumulated value is used for calculation The normalized denominator of .
[0049] If the word Located inside state q, by calculating Get sending probability In this way, the dependency between a word and the previous word is taken into account, reflecting the arrangement characteristics of the words within the state.
[0050] In order to correct the influence of the SVM classification score on the emission probability, the Sigmoid function is introduced to map the SVM classification score v(σ) into a correction factor. The formula is: ,in,a and b Determined by 4-fold cross validation, for example a =−3.9, b =0.21 to ensure that the corrected emission probability can effectively improve the confidence of the SVM and make it reach or exceed 0.90 on average.
[0051] In this way, the candidate metadata fields obtained by SVM classification can be extracted and corrected more finely, thereby achieving more accurate field boundaries and internal structures, and eliminating errors caused by fuzzy boundaries or semantic ambiguity, thereby ensuring that the target metadata finally obtained achieves the expected accuracy and consistency.
[0052] Furthermore, target metadata for the document data to be extracted is obtained based on the candidate metadata, including: obtaining a first similarity between the candidate metadata and a preset standard field; obtaining a second similarity between the candidate text block data and the historical text block data; and obtaining a contextual matching score for the candidate text block data. The candidate text block data is the text block data corresponding to the candidate metadata. The historical text block data is the text block data corresponding to metadata in the historical document data. The target metadata is obtained based on the first similarity, the second similarity, and the contextual matching score.
[0053] Furthermore, obtaining the first similarity between the candidate metadata and the preset standard field includes: calculating Get the first similarity, where is the first similarity between the candidate metadata and the preset standard field, A is the candidate metadata vector, and B is the standard field vector. . is the candidate metadata vector A p Components, B p is the first standard field vector B p components. p is the dimension number of the candidate metadata vector and the standard field vector.
[0054] Optionally, the preset standard field is obtained by obtaining the industry field to which the document data corresponding to the text block data belongs. Using the industry field, a table lookup operation is performed in a preset second data table to find the standard field corresponding to the industry field. The second data table stores the correspondence between the industry field and the standard field.
[0055] Optionally, obtaining the second similarity between the candidate text block data and the historical text block data includes: calculating Get the second similarity, where is the second similarity between the candidate text block data and the historical text block data, C is the vector representation of the candidate text block data, and D is the vector representation of the historical text block data.
[0056] Optionally, the second similarity is obtained using edit distance or Jaccard similarity. For example, by calculating Get the second similarity, where is the candidate text block data, is the historical text block data, and d is the edit distance.
[0057] Furthermore, obtaining a contextual matching score for the candidate text block data includes obtaining a degree of matching between the candidate text block data and a preset structural anchor point, and obtaining a cosine similarity between the candidate text block data and its adjacent text block data. The contextual matching score for the candidate text block data is obtained based on the degree of matching between the candidate text block data and the preset structural anchor point, and the cosine similarity between the candidate text block data and its adjacent text block data.
[0058] Furthermore, obtaining the matching degree between the candidate text block data and the preset structural anchor point includes: calculating Obtain the matching degree between the candidate text block data and the preset structural anchor point. is the matching degree between the candidate text block data and the p-th structural anchor point, is the parameter that controls the decay rate. is the distance between the candidate text block data and the pth structural anchor point. Preset structural anchor points include titles, tables, or chapter separators.
[0059] Furthermore, according to the matching degree between the candidate text block data and the preset structural anchor point and the cosine similarity between the candidate text block data and its adjacent text block data, the context matching score of the candidate text block data is obtained, including: Get the context matching score of the candidate text block data, is the context matching score for the candidate text block data, is the matching degree between the candidate text block data and the title structure anchor point, is the matching degree between the candidate text block data and the table structure anchor point, is the cosine similarity between the candidate text block data and its adjacent text block data. is the preset first weight coefficient, is the preset first weight coefficient, is the preset first weight coefficient, and + + = 1.
[0060] Furthermore, obtaining target metadata based on the first similarity, the second similarity, and the context match score includes obtaining a confidence level for the candidate metadata based on the first similarity, the second similarity, and the context match score. If the confidence level is greater than or equal to a preset threshold, the corresponding candidate metadata is determined as the target metadata. For example, the preset threshold is 0.9. In this case, due to the high confidence level of the candidate metadata, the candidate metadata is directly determined as the target metadata, thereby ensuring that the obtained target metadata has a high accuracy rate.
[0061] When the confidence of the candidate metadata is less than a preset threshold, the preset pre-trained large language model is used to obtain the target metadata.
[0062] Specifically, a pre-trained large language model is used to obtain target metadata. This includes inputting candidate metadata corresponding to the confidence level, the corresponding candidate text block data, structural information about the candidate text block data in the corresponding document data, adjacent text information about the candidate text block data, a first similarity between the candidate metadata and a preset standard field, a second similarity between the candidate text block data and historical text block data, and a contextual matching score for the candidate text block data into the pre-trained large language model to obtain the target metadata. In this way, if the confidence level of the candidate metadata is low, the large language model uses the contextual information and domain knowledge of the text block data to conduct a deep semantic understanding of the candidate metadata, correct ambiguous or fuzzy classification results, and ensure that the output target metadata matches the actual semantics. This further improves the accuracy of the target metadata. For example, for the candidate metadata "circuit breaker tripped," the large language model can confirm the semantic correspondence between it and the equipment trip fault in the power equipment, propose adjustment suggestions, and output the target metadata as "equipment trip fault."
[0063] In one embodiment, the structural information of the candidate text block data in the corresponding document data includes one or more of the following: page number, title, paragraph segmentation, table location, and header and footer information. The adjacent text information of the candidate text block data includes a preset number of lines of text information preceding the candidate text block data and a preset number of lines of text information following the candidate text block data. For example, the preset number of lines is 3. Furthermore, based on the first similarity, the second similarity, and the contextual matching score, the confidence level of the candidate metadata is obtained, including: By calculation Get the confidence of the candidate metadata. is the confidence of the candidate metadata, is the first weight parameter, is the second weight parameter, is the third weight parameter, and .For example, , , 0.3.
[0064] Optionally, after obtaining the target metadata, the method further includes: performing dynamic semantic mapping on the target metadata using a preset ontology. Dynamic semantic mapping is performed on each target metadata x, that is, selecting a target node y that best matches x in the ontology, thereby achieving standardized mapping.
[0065] Using text vectorization and cosine similarity to calculate the target text block data, target metadata, the adjacent text information of each preset number of lines before and after the text block data, and the structural information of the target text block data in the document data, the candidate node set Y that meets the preset conditions of semantic similarity and context matching with the target metadata x is selected from the ontology. = { y 1 , y 2 , …, y n The structural information of the target text block data in the document data includes: header, footer, title or table position, etc.
[0066] By calculation Obtain the posterior probability between each candidate node and the target metadata. is the probability of observing the xth target metadata under the condition of the vth candidate node, is the prior probability of the vth candidate node, Represents the g-th semantic grouping under a given concept y The conditional probability of the xth target metadata element appearing under the condition is used to quantify the association strength between x and the semantic grouping.
[0067] Define the loss function , when the mapping is correct, the loss value When the mapping is wrong, the loss value . Then the mapping risk function is: .in, and .when When , the risk value is simplified to
[0068] The mapping solution with the lowest risk value is selected from all candidate mappings. For each target metadata x, in the candidate set Y, the candidate node y with the lowest risk value R(x, y) and a mapping confidence P(y|x) that meets a preset threshold, for example, P(y|x) ≥ 0.96, is selected as the final mapping result. Ensure that the average mapping risk is less than 0.05 and the mapping accuracy reaches or exceeds 96%. For example, accurately map the extracted "circuit breaker trip" to the "equipment trip fault" node in the ontology.
[0069] In some embodiments, the preset ontology is a power industry domain ontology. The preset ontology is constructed by utilizing multi-source, multi-dimensional historical data such as grid operation specifications, equipment maintenance standards, and historical data, such as grid operation monitoring, equipment inspection and maintenance records, fault alarm and processing logs, power outage recovery events, equipment performance aging statistics, and topology change and load migration records, to construct a power industry domain ontology. That is, within the power industry, a knowledge model that formalizes and structures the core concepts, entities, and their interrelationships in various aspects such as grid operation and equipment maintenance. This ontology formalizes core concepts such as grid zoning, substation equipment, fault types, maintenance processes, and regional divisions, and clarifies the hierarchical structure and association relationships between these concepts. This ontology includes core concepts such as grid zoning, substation equipment, fault types, maintenance processes, and regional divisions.
[0070] The ontology is described using the OWL language, clarifying the hierarchical relationships and attribute constraints between nodes. For example, "circuit breaker trip" is defined as an instance of the "equipment failure" category and associated with attributes such as "downtime" and "maintenance record." During design, ensure that the ontology coverage reaches above 95%, providing an accurate and structured reference standard for subsequent semantic mapping.
[0071] Optionally, after obtaining the target metadata, the method further includes: obtaining display information, and displaying the display information via a display device. The display information includes the target metadata and the confidence level corresponding to the target metadata. The display information also includes the mapped candidate nodes and their risk values. For example, the target metadata, confidence level, mapped candidate nodes, and risk values corresponding to each text block data are displayed in real time via a web-based graphical interface. This enables intuitive presentation to experts in the power industry. This facilitates manual review and modification of low-confidence or high-risk target metadata and mapping results by experts. Feedback information is recorded in a database via a system interface and serves as a data basis for model improvement. Based on expert feedback, the system regularly triggers a parameter retraining module to iteratively optimize the SVM classifier, improved BiHMM model, large language model-assisted semantic processing module, and semantic mapping algorithm, forming a closed-loop feedback mechanism to ensure that the system's extraction accuracy remains above 90% and the mapping accuracy remains above 96% during actual operation.
[0072] Combine Figure 2 As shown, an embodiment of the present disclosure provides an apparatus 200 for metadata extraction, comprising a document data acquisition module 201, a text block data acquisition module 202, a feature vector acquisition module 203, a candidate metadata acquisition module 204, and a target metadata acquisition module 205. The document data acquisition module 201 is configured to acquire document data to be extracted. The text block data acquisition module 202 is configured to acquire text block data of the document data to be extracted. The feature vector acquisition module 203 is configured to acquire feature vectors of the text block data. The candidate metadata acquisition module 204 is configured to acquire candidate metadata based on the feature vectors. The target metadata acquisition module 205 is configured to acquire target metadata for the document data to be extracted based on the candidate metadata.
[0073] The device for metadata extraction provided by the embodiments of the present disclosure can obtain feature vectors of text block data from the document data to be extracted, first obtain candidate metadata based on the feature vectors of the text block data, and then obtain target metadata based on the candidate metadata. Obtaining candidate metadata based on the feature vectors of the text block data from the document data to be extracted can adapt to metadata extraction from documents with different formats and inconsistent professional terminology, and further obtain target metadata based on the candidate metadata, thereby achieving higher accuracy in the obtained target metadata.
[0074] Optionally, the apparatus for metadata extraction further includes a display module, wherein the display module is configured to obtain display information and display the display information.
[0075] Combine Figure 3 As shown, an embodiment of the present disclosure provides an electronic device 300, including a processor 304 and a memory 301 storing program instructions. Optionally, the electronic device may also include a communication interface 302 and a bus 303. The processor 304, communication interface 302, and memory 301 may communicate with each other via bus 303. The communication interface 302 may be used for information transmission. The processor 304 may invoke the program instructions in the memory 301 to execute the metadata extraction method of the above embodiment.
[0076] In addition, the logic instructions in the memory 301 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0077] Memory 301, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions / modules corresponding to the methods in the embodiments of the present disclosure. Processor 304 executes the program instructions / modules stored in memory 301 to perform functional applications and data processing, thereby implementing the metadata extraction methods in the above-described embodiments.
[0078] The memory 301 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal device. Furthermore, the memory 301 may include high-speed random access memory and non-volatile memory.
[0079] An embodiment of the present disclosure provides a storage medium storing program instructions, which are executed by a processor to implement the above-mentioned method for metadata extraction.
[0080] The technical solutions of the embodiments of the present disclosure may be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present disclosure. The aforementioned storage medium may be a non-transitory storage medium, including: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, and other media that can store program code, or a transient storage medium.
[0081] The above description and the accompanying drawings sufficiently illustrate the embodiments of the present disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible variations. Unless expressly required, individual components and functions are optional, and the order of operations may vary. Portions and features of some embodiments may be included in or replace portions and features of other embodiments. Moreover, the terms used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and claims, the singular forms "a", "an" and "the" are intended to also include the plural forms unless the context clearly indicates otherwise. Similarly, the term "and / or" as used in this application means any and all possible combinations of one or more of the associated listings. In addition, when used in this application, the term "comprise" and its variations "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, an element defined by the sentence "comprising a..." does not exclude the presence of other identical elements in the process, method or device that includes the element. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments can be referenced to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can be referred to the description of the method part.
[0082] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0083] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices and equipment) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units may be merely a logical functional division. In actual implementation, other divisions may be used, such as combining or integrating multiple units or components into another system, or omitting or disabling some features. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be through interfaces, indirect couplings or communication connections between devices or units, and may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of these units may be selected to implement the embodiments according to actual needs. Furthermore, the functional units in the disclosed embodiments may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0084] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architectures, functions and operations of the systems, methods and computer program products according to the embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the module, program segment or part of the code contains one or more executable instructions for implementing the specified logical functions. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in an order different from that disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action, or may be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A method for metadata extraction, characterized in that: include: Get the document data to be extracted; Acquire text block data of the document data to be extracted; Obtaining a feature vector of the text block data; Obtaining candidate metadata according to the feature vector; According to the candidate metadata, target metadata of the document data to be extracted is acquired.
2. The method according to claim 1, characterized in that Acquiring the text block data of the document data to be extracted includes: Extracting text information from the document data to be extracted; The text information is divided into text blocks to obtain text block data.
3. The method according to claim 1, characterized in that Obtaining a feature vector of the text block data includes: Obtaining context features of the text block data, obtaining a matching score between a preset keyword and the text block data, obtaining structural features of the text block data, and obtaining a word frequency and a TF-IDF value of each word in the text block data; A feature vector of the text block data is obtained according to the context feature, the matching score, the structural feature, each word frequency and each TF-IDF value.
4. The method according to claim 1, wherein According to the feature vector, candidate metadata is obtained, including: The feature vector is input into a preset classification model to obtain candidate metadata.
5. The method according to claim 1, wherein Acquiring target metadata of the document data to be extracted based on the candidate metadata includes: Obtaining a first similarity between the candidate metadata and a preset standard field, obtaining a second similarity between the candidate text block data and the historical text block data, and obtaining a context matching score for the candidate text block data; wherein the candidate text block data is the text block data corresponding to the candidate metadata, and the historical text block data is the text block data corresponding to the metadata in the historical document data; Target metadata is obtained according to the first similarity, the second similarity, and the context matching score.
6. The method according to claim 5, characterized in that According to the first similarity, the second similarity and the context matching score, target metadata is obtained, including: Obtaining a confidence level of the candidate metadata according to the first similarity, the second similarity, and the context matching score; When the confidence level is greater than or equal to a preset threshold, determining the corresponding candidate metadata as target metadata; When the confidence level is less than a preset threshold, the target metadata is obtained using a preset pre-trained large language model.
7. The method according to any one of claims 1 to 6, characterized in that After obtaining the target metadata, it also includes: Dynamic semantic mapping is performed on the target metadata using a preset ontology.
8. A device for metadata extraction, characterized in that: include: A document data acquisition module is configured to acquire document data to be extracted; A text block data acquisition module is configured to acquire the text block data of the document data to be extracted; A feature vector acquisition module is configured to acquire a feature vector of the text block data; a candidate metadata acquisition module, configured to acquire candidate metadata according to the feature vector; The target metadata acquisition module is configured to acquire target metadata of the document data to be extracted based on the candidate metadata.
9. An electronic device comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the method for metadata extraction according to any one of claims 1 to 7 when running the program instructions.
10. A storage medium storing program instructions, characterized in that: The program instructions are executed by a processor to implement the method for metadata extraction according to any one of claims 1 to 7.