Structured Information Extraction Method and Device Based on Multivariate Annotation Strategy

Through the structured information extraction method based on the multi-dimensional annotation strategy, the problem of the inability to correspond one by one to one of the timelines and information of scholars' education or work experience in the prior art is solved, and the accurate structured extraction of scholars' resume information is achieved.

CN113836891BActive Publication Date: 2025-06-13BEIJING ZHIPU XINGYAO TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111016304.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-31
Publication Date
2025-06-13
Estimated Expiration
2041-08-31

AI Technical Summary

Technical Problem

The prior art is difficult to automatically, accurately and quickly extract structured information of scholars' education or work experience from massive scattered and unstructured data, especially the timeline and experience information cannot be matched one by one.

Method used

The structured information extraction method based on the multivariate annotation strategy is adopted to realize the structured extraction of scholars' resume information by crawling the scholar's homepage, cleaning and sentence processing, regular expression matching date, short text classification model filtering resume text, multi-label sequence annotation, BERT-Bi-LSTM-CRF model training and prediction.

Benefits of technology

The scholar's school, major, degree and other information are successfully matched with the start and end time, and accurate and structured data are obtained, which meets the application needs of actual scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113836891B_ABST
    Figure CN113836891B_ABST
Patent Text Reader

Abstract

This application proposes a structured information extraction method and device based on a multi - annotation strategy. The method includes: crawling the homepage of scholars, and performing cleaning processing and sentence - splitting processing on the homepage; matching sentences through regular expressions and unifying the dates in different formats in the sentences; screening out the texts containing scholars' resumes through a preset short - text classification model; performing multi - label sequence annotation on the texts based on the multi - annotation strategy, and splitting the obtained label dataset into a training set, a validation set, and a test set. Training a BERT - Bi - LSTM - CRF model based on the data in the training set; predicting the results of multi - label sequence annotation through the trained model, and evaluating the prediction effect. This application regards the task of structured information extraction as a multi - label sequence annotation task, and combines a deep - learning network model to perform structured extraction on information such as scholars' resumes, corresponding information such as schools attended, majors, degrees, etc. with the timeline one by one to obtain accurate and structured data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information extraction, and in particular, to a structured information extraction method and device based on a multi-labeling strategy. Background Art

[0002] Currently, there are hundreds of millions of experts and scholars globally, and most of the data of these experts and scholars is presented on the Internet in heterogeneous and unstructured forms. This data contains a large amount of valuable information, such as basic information (email, title, work unit, etc.), educational experience (institutions attended, academic degrees, etc.), and work resumes (work units, titles, etc.). Due to the diverse data sources and unstructured storage, it is difficult to directly construct a multi-precision talent semantic portrait of a scholar to meet the intelligent talent analysis needs in various different scenarios and data dimensions. Therefore, how to automatically, accurately, and quickly extract valuable information from a large amount of scattered and unstructured data and sort out the relevant educational or work experiences of experts and scholars has become a hot issue of concern in the academic and industrial circles.

[0003] In related technologies, the extraction of educational or work experiences mainly uses rule / regular expression-based text matching methods or traditional sequence labeling methods. However, these studies do not consider corresponding the timeline and experience information one by one. For example, in the case of educational experience, information such as institutions attended, majors, degrees, start and end times, etc. are not corresponded one by one, which cannot meet the requirements of actual application scenarios. Summary of the Invention

[0004] This application aims to solve at least one of the technical problems in the related technologies to some extent.

[0005] To this end, the first object of this application is to propose a structured information extraction method based on a multi-labeling strategy. The method first crawls the homepage of a scholar, and performs cleaning and sentence splitting processing on the homepage; matches and converts dates in different formats in the sentence through regular expressions; filters out the text representing the resume of the scholar through a preset short text classification model; performs multi-label sequence labeling on the text based on the multi-labeling strategy to obtain a label data set, and divides the label data set into a training set, a validation set, and a test set; trains a BERT-Bi-LSTM-CRF model based on the data in the training set; predicts the results of multi-label sequence labeling through the trained model. This method combines the multi-labeling strategy and a deep learning network model to extract the resume information of scholars, solves the problem of resume structuring of scholars, and can correspond information such as institutions attended, majors, degrees, etc. of scholars with start and end times to obtain accurate and structured data, meeting the application requirements of actual scenarios.

[0006] The second object of this application is to propose a structured information extraction device based on a multi-labeling strategy.

[0007] The third object of this application is to propose a non - temporary computer - readable storage medium.

[0008] To achieve the above object, an embodiment of the first aspect of this application proposes a structured information extraction method based on a multi - label annotation strategy, including the following steps:

[0009] Crawl the homepage of the scholar, and perform cleaning processing and sentence - splitting processing on the homepage;

[0010] Match each sentence through regular expressions to obtain dates in different formats on the homepage, and convert the format of each date into a preset date format;

[0011] Classify each sentence through a preset short - text classification model, and screen out the text containing the scholar's resume;

[0012] Perform sequence annotation on the text based on a multi - label annotation strategy to obtain a label data set, and divide the label data set into a training set, a validation set, and a test set according to a preset ratio;

[0013] Train a Bidirectional Encoder Representations from Transformers - Bidirectional Long Short - Term Memory Network - Conditional Random Field (BERT - Bi - LSTM - CRF) model based on the data in the training set;

[0014] Predict the results of multi - label sequence annotation through the trained BERT - Bi - LSTM - CRF model, and evaluate the prediction effect of the BERT - Bi - LSTM - CRF model.

[0015] Optionally, in an embodiment of this application, the cleaning processing of the homepage includes: converting the page language of the homepage into a lightweight markup language Markdown; converting non - ASCII characters in the homepage into Unicode characters and deleting illegal characters; converting the font and case of each character in the homepage; the sentence - splitting processing of the homepage includes: dividing the content in the homepage into multiple sentences through a preset toolkit.

[0016] Optionally, in an embodiment of this application, the short - text classification model includes a vector representation layer, a convolutional layer, a max - pooling layer, and a fully - connected layer. The classifying each sentence through the preset short - text classification model includes: mapping each character in each sentence into a vector of a preset dimension; performing convolutional processing on each sentence through different convolutional kernel sizes to obtain multiple one - dimensional column vectors; taking the maximum value in each one - dimensional column vector and splicing each maximum value to obtain a pooled vector; classifying the pooled vector through a softmax function to determine the category to which each sentence belongs.

[0017] Optionally, in an embodiment of the present application, the sequence labeling of the text based on the multi-source annotation strategy includes: performing position part labeling on each word in the text through BIO sequence labeling; performing entity type part labeling on each word in the text; setting a number for each piece of experience, and performing experience stage part labeling on each word in the text.

[0018] Optionally, in an embodiment of the present application, training a BERT-Bi-LSTM-CRF model based on the data in the training set includes: converting each sentence in the training set into a word-level sequence, and inputting each word-level sequence into the BERT model to generate a vector of a first size based on context information; inputting the vector of the first size into a bidirectional long short-term memory network Bi-LSTM for feature extraction to output a vector of a second size; inputting the vector of the second size into a conditional random field CRF model for decoding after passing through a fully connected layer, training the model based on the scores of the label sequences corresponding to the word-level sequences output by the CRF model, and calculating the target annotation sequence.

[0019] Optionally, in an embodiment of the present application, the CRF model calculates the score of the label sequence corresponding to the word-level sequence through the following formula:

[0020]

[0021] where X represents the word-level sequence, y represents the label sequence predicted by the CRF model for the word-level sequence, P i,j represents the probability that the i-th word in the input word-level sequence corresponds to the j-th label in the label data set, and T i,j represents the transition probability from label i to label j for consecutive words in the word-level sequence.

[0022] Optionally, in an embodiment of the present application, after obtaining the score of the label sequence corresponding to the word-level sequence, a loss function is defined through the maximum log-likelihood function, and the BERT-Bi-LSTM-CRF model is trained using the gradient descent method. The maximum log-likelihood function is expressed as follows:

[0023]

[0024] where p(y|X) represents the probability value that each label sequence is the correct label sequence, and Y x represents all possible label sequences;

[0025] The target annotation sequence is calculated through the following formula:

[0026]

[0027] Among them, y * is the output sequence with the maximum conditional probability.

[0028] Optionally, in an embodiment of the present application, evaluating the prediction effect of the BERT-Bi-LSTM-CRF model includes:

[0029] Calculating the precision rate, recall rate, and comprehensive evaluation value of the results predicted by the fine-tuned BERT-Bi-LSTM-CRF model;

[0030] Evaluating the fine-tuned generative training model according to the precision rate, the recall rate, and the comprehensive evaluation value; wherein, the precision rate, the recall rate, and the comprehensive evaluation value are calculated by the following formulas:

[0031]

[0032]

[0033] Among them,

[0034] Among them, P is the precision rate, R is the recall rate, F1 is the comprehensive evaluation value, m is the number of extracted records, n is the number of labeled records, and k is the number of elements of record i in the labeled data.

[0035] To achieve the above object, the second aspect embodiment of the present application proposes a structured information extraction device based on a multi-labeling strategy of the present invention, including the following modules:

[0036] A homepage generation module, configured to obtain the homepage of a scholar, and perform cleaning processing and sentence splitting processing on the homepage;

[0037] A date generation module, configured to match each sentence through a regular expression to obtain dates in different formats in the homepage, and convert the format of each date into a preset date format;

[0038] A sentence classification module, configured to classify each sentence through a preset short text classification model, and screen out the text representing the resume of the scholar;

[0039] A multi-labeling module, configured to perform sequence labeling on the text based on a multi-labeling strategy to obtain a label data set, and divide the label data set into a training set, a validation set, and a test set according to a preset ratio;

[0040] A model training module, configured to train a Bidirectional Encoder Representations from Transformers - Bidirectional Long Short - Term Memory Network - Conditional Random Field (BERT - Bi - LSTM - CRF) model based on the data in the training set;

[0041] A prediction module, configured to predict the results of multi - label sequence annotation through the trained BERT - Bi - LSTM - CRF model, and evaluate the prediction effect of the BERT - Bi - LSTM - CRF model.

[0042] Optionally, in an embodiment of the present application, the homepage generation module is specifically configured to: convert the page language of the homepage into a lightweight markup language Markdown; convert non - ASCII characters in the homepage into Unicode characters and delete illegal characters; convert the font and case of each character in the homepage; and perform sentence splitting on the homepage, including: dividing the content in the homepage into multiple sentences through a preset toolkit.

[0043] Optionally, in an embodiment of the present application, the sentence classification module is specifically configured to: map each character in each sentence into a vector of a preset dimension; perform convolution processing on each sentence through different convolution kernel sizes to obtain multiple one - dimensional column vectors; take the maximum value in each one - dimensional column vector and splice each maximum value to obtain a pooled vector; and classify the pooled vector through a softmax function to determine the category to which each sentence belongs.

[0044] Optionally, in an embodiment of the present application, the multi - element annotation module is specifically configured to: perform position part annotation on each character in the text through BIO sequence annotation; perform entity type part annotation on each character in the text; set a number for each experience segment, and perform experience stage part annotation on each character in the text.

[0045] Optionally, in an embodiment of the present application, the model training module is specifically configured to: convert each sentence in the training set into a word - level sequence, and input each word - level sequence into a BERT model to generate a vector of a first size based on context information; input the vector of the first size into a Bidirectional Long Short - Term Memory Network (Bi - LSTM) for feature extraction and output a vector of a second size; input the vector of the second size into a Conditional Random Field (CRF) model after passing through a fully - connected layer, perform model training based on the scores of the label sequences corresponding to the word - level sequences output by the CRF model, and calculate the target annotation sequence.

[0046] Optionally, in an embodiment of the present application, the CRF model calculates the score of the tag sequence corresponding to the word-level sequence through the following formula:

[0047]

[0048] where X represents the word-level sequence, y represents the tag sequence after the CRF model predicts the word-level sequence, and P i,j represents the probability that the i-th word in the input word-level sequence corresponds to the j-th tag in the tag dataset, and T i,j represents the transition probability of consecutive words in the word-level sequence from tag i to tag j.

[0049] Optionally, in an embodiment of the present application, the model training module is specifically configured to: after obtaining the score of the tag sequence corresponding to the word-level sequence, define a loss function through the maximum log-likelihood function, and use the gradient descent method to train the BERT-Bi-LSTM-CRF model. The maximum log-likelihood function is expressed as follows:

[0050]

[0051] where p(y|X) represents the probability value that each tag sequence is the correct tag sequence, and Y x represents all possible tag sequences;

[0052] Calculate the target annotation sequence through the following formula:

[0053]

[0054] where y * is the output sequence with the maximum conditional probability.

[0055] Optionally, in an embodiment of the present application, the prediction module is specifically configured to:

[0056] Calculate the precision rate, recall rate, and comprehensive evaluation value of the results predicted by the trained BERT-Bi-LSTM-CRF model;

[0057] Evaluate the fine-tuned generative training model according to the precision rate, the recall rate, and the comprehensive evaluation value; where the precision rate, the recall rate, and the comprehensive evaluation value are calculated through the following formula:

[0058]

[0059]

[0060] where

[0061] Among them, P is the precision rate, R is the recall rate, F1 is the comprehensive evaluation value, m is the number of extracted records, n is the number of labeled records, and k is the number of elements of record i in the labeled data.

[0062] The technical solution provided by the embodiment of the present application at least brings the following beneficial effects: The method first crawls the homepage of the scholar, and performs cleaning processing and sentence segmentation processing on the homepage; matches and converts dates in different formats in the sentence through regular expressions; filters out texts representing the educational or work experience of the scholar through a preset short text classification model; performs sequence labeling on the text based on a multi-labeling strategy to obtain a label data set, and divides the label data set into a training set, a validation set, and a test set; trains a Bidirectional Encoder Representations from Transformers-Bidirectional Long Short-Term Memory-Conditional Random Field (BERT-Bi-LSTM-CRF) model based on the data in the training set; predicts the results of multi-label sequence labeling through the trained model. This method combines a multi-labeling strategy and a deep learning network model to perform structured extraction of the resume information of scholars. It can not only extract the basic information of scholars, but also solve the problem that the resume information such as the educational or work experience of scholars does not match the timeline, and makes a one-to-one correspondence between information such as the institutions attended, majors, and degrees of scholars and the start and end times, obtaining accurate and structured data, meeting the application requirements of the actual scenario.

[0063] To implement the above embodiment, the third aspect of the present application also proposes a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the structured information extraction method and device based on the multi-labeling strategy in the above embodiment.

[0064] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0065] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:

[0066] Figure 1 is a flowchart of a structured information extraction method based on a multi-labeling strategy proposed by an embodiment of the present application;

[0067] Figure 2 is a structural diagram of a TextCNN network proposed by an embodiment of the present application;

[0068] Figure 3 is a structural diagram of a structured information extraction model based on BERT-Bi-LSTM-CRF proposed by an embodiment of the present application;

[0069] Figure 4 It is a schematic structural diagram of a BERT proposed in an embodiment of the present application;

[0070] Figure 5 It is a schematic structural diagram of an LSTM proposed in an embodiment of the present application;

[0071] Figure 6 It is a schematic flowchart of a specific structured information extraction method based on a multi - annotation strategy proposed in an embodiment of the present application;

[0072] Figure 7 It is a schematic structural diagram of a structured information extraction device based on a multi - annotation strategy proposed in an embodiment of the present application. Detailed implementation manners

[0073] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0074] The method and device for structured information extraction based on a multi - annotation strategy proposed in an embodiment of the present invention will be described below with reference to the accompanying drawings.

[0075] Figure 1 It is a flowchart of a structured information extraction method based on a multi - annotation strategy proposed in an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0076] Step 101, crawl the homepage of the scholar, and perform cleaning processing and sentence - splitting processing on the homepage.

[0077] Among them, the homepage of the scholar includes various basic information of the scholar, such as gender, date of birth, research direction, unit, professional title, position, work experience, educational background, etc., as well as information such as the research achievements of the scholar. In the embodiment of the present application, the homepage of the scholar can be obtained through relevant crawling codes or crawling tools and other crawling methods.

[0078] In the technical solution of the present application, the crawling of the homepage of the scholar complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0079] In one embodiment of the present application, the homepage is cleaned, including: converting the page language of the homepage into a lightweight markup language Markdown; converting non-ASCII characters in the homepage into Unicode characters and deleting illegal characters; converting the font and uppercase of each character in the homepage. The homepage is sentence-segmented, including: dividing the content in the homepage into multiple sentences through a preset toolkit.

[0080] Specifically, after obtaining the scholar's homepage, the obtained homepage is cleaned. In one embodiment of the present application, the cleaning step can be to first convert the page language of the homepage into a lightweight markup language Markdown. Since the crawled homepage is in html format, in order to facilitate subsequent operations, the html2text toolkit is used to remove various useless tags in the html file to form a plain text page containing only actual content. Then, the characters are processed, and the non-ASCII characters in the homepage text are converted into Unicode characters, and illegal characters are deleted. Then, the font of each character in the homepage text is converted from traditional to simplified, full and half-width, and the font format and font size are unified. The last step is to perform sentence processing on the homepage, and the text in the homepage can be divided into multiple sentences by using a sentence toolkit such as the SentenceSpliter package of the preset Pyltp tool.

[0081] Step 102, matching each sentence with a regular expression to obtain dates in different formats on the homepage, and converting the format of each date into a preset date format.

[0082] It can be understood that a scholar's homepage may contain sentences expressing dates in different formats. In order to facilitate the subsequent matching of the scholar's education or work experience and other information to be marked with the date, it is necessary to first unify the date format.

[0083] In specific implementation, in one embodiment of the present application, a preset regular expression can be used to match each sentence after the sentence segmentation, and the string matching pattern described by the regular expression can be used to determine whether each sentence contains a date, and the date can be taken out from each sentence, and then the regular expression can be used to replace the year and month in the sentence with a preset format.

[0084] As an example, the specific regular expression is as follows:

[0085] (\s{0,2}

[12] \d{3})? [year. / -]? (\d{1,2})? [month]? (\s{0,2}[-to–—~-~_,\t]{1,2}\s{0,2})? (

[12] \d{3})[year. / -]? (\d{1,2})? [month]? ".

[0086] Through this regular expression, the years and months extracted from each clause are uniformly converted into the format of YYYY or YYYY - MM.

[0087] Step 103: Classify each sentence through a preset short - text classification model, and filter out the text containing the resume of the scholar.

[0088] Among them, the resume information of the scholar is the information to be sequence - labeled, which can be the scholar's date of birth, professional title, employment unit, educational experience, work experience, etc., which can be represented in a structured way. Before performing multi - label annotation to extract the structured information of the scholar, this application first determines the text containing the resume information of the scholar.

[0089] In an embodiment of this application, a short - text classification model can be first constructed based on the text convolutional neural network TextCNN. As Figure 2 shown, the preset short - text classification model of this application includes a vector representation layer, a convolutional layer, a max - pooling layer, and a fully - connected layer. Classify each sentence through this short - text classification model, and filter out the text representing the resume information. The specific steps are as follows:

[0090] First, for the convenience of subsequent processing, through the vector representation layer, natural language is digitized, that is, each word in each sentence is mapped into a vector of a preset dimension. In an embodiment of this application, each word is mapped into a 100 - dimensional vector through word2vec. Then, in the convolutional layer, each sentence is convolved through different convolutional kernel sizes kernel_size, and the sentence is converted into a one - dimensional vector through convolution processing to obtain multiple one - dimensional column vectors. In the max - pooling layer, the maximum value is taken from each one - dimensional vector, and each maximum value is concatenated to obtain the pooled vector. Finally, in the fully - connected layer, the pooled vector is classified through the softmax function, and 5 dimensions are output, namely basic information, educational experience, work resume, project information, and others. These five categories are used to determine the category to which each sentence belongs. In addition, to prevent the model from overfitting, this application also introduces L2 regularization processing and dropout processing in the model to perform weight decay and randomly deactivate some neural units in the current layer of TextCNN.

[0091] Step 104: Perform sequence annotation on the text based on a multi - label annotation strategy to obtain a label dataset, and divide the label dataset into a training set, a validation set, and a test set according to a preset ratio.

[0092] It should be noted that since there may be multiple experiences of a scholar in one sentence, some fields are shared among multiple experiences. Therefore, when this application performs sequence labeling on the classified sentences, a multi-labeling method is adopted. As an example, this application adopts a three-way labeling method, that is, each word in the sentence is marked through a three-way labeling strategy, so that each word in the sentence corresponds to a unique tag, in order to distinguish the same characters in other experiences and facilitate obtaining the corresponding relationship of experience information.

[0093] In an embodiment of this application, multi-label sequence labeling is performed on the text based on the multi-labeling strategy, including: performing position part labeling on each word in the text through BIO sequence labeling; performing entity type part labeling on each word in the text; setting a number for each experience, and performing experience stage part labeling on each word in the text.

[0094] Specifically, the multi-labeling strategy in the embodiment of this application consists of the following three parts:

[0095] Position part: Use the traditional BIO sequence labeling method to encode the position information of each word in an entity, where B represents that the word belongs to the start of an entity, I represents that the word is in a non-start position of the entity, that is, the middle position, and O represents that the word does not belong to any part of the entity. The entity here is the entity type in the information to be labeled such as the learning or working experience of the scholar, such as the work unit or school, etc.

[0096] Entity type part: Associate each word with the type information of the entity. The type information of the entity includes: "Univ", "Deg", "Major", etc., which represent the school attended, degree, and major respectively.

[0097] Experience stage part: First, set a number for each experience of the scholar, such as 1, 2,..., etc. 1 represents the first experience, 2 represents the second experience, and so on. In addition, in order to be able to extract information independent of the timeline and enrich the result of information extraction, this application also sets an experience stage flag numbered 0. For example, the information in this section of the experience includes the date of birth, etc.

[0098] Thus, this application sets a multi-labeling strategy, gives all the corresponding tags for each word's token according to each experience, realizes sequence labeling of the text, and then combines each tag obtained after labeling to generate a tag data set.

[0099] Optionally, in an embodiment of this application, the data set with the multi-labeling strategy set is split into a training set, a validation set, and a test set according to the ratio of 7:1.5:1.5, which is convenient for subsequent training of the model and verification of the prediction effect of the model, etc.

[0100] Step 105: Train a Bidirectional Encoder Representations from Transformers - Bidirectional Long Short - Term Memory Network - Conditional Random Field (BERT - Bi - LSTM - CRF) model based on the data in the training set.

[0101] It should be noted that after the text is annotated with the above BIO annotation mechanism, the obtained labels are not independent of each other, and it is impossible to correspond information such as the work or educational experience of scholars in the text according to these labels. Therefore, after obtaining the label dataset, the present application processes the relationship between the labels through the BERT - Bi - LSTM - CRF model.

[0102] Among them, the present application pre - constructs a Bidirectional Encoder Representations from Transformers (abbreviated as BERT) - Bidirectional Long Short - Term Memory Network (Bi - LSTM) - Conditional Random Field (CRF) model, that is, the BERT - Bi - LSTM - CRF model. By making this model perform sequence annotation and calculating the optimal annotation sequence, that is, determining the most matching label in the label dataset corresponding to each word through this model, the model is trained. After the training is completed, the optimal label sequence corresponding to the word sequence to be labeled is output through this model, that is, the result of predicting the multi - label sequence annotation, and structured data corresponding to each experience of the scholar can be obtained, realizing the structured extraction of the scholar's resume information.

[0103] In an embodiment of the present application, the structure of the BERT - Bi - LSTM - CRF model is as Figure 3 shown. When performing sequence annotation through this model, the following steps are included:

[0104] First, convert each sentence in the training set into a word - level sequence, obtain the token_id according to the word - level sequence using the word piece tokenizer, and input the token_id of each word - level sequence into the BERT model to generate a vector of the first dimension based on context information. The first dimension is batch_size * max_seq_len * emb_size.

[0105] Then, input the sequence output by the BERT model, that is, each vector of the first dimension, into the Bidirectional Long Short - Term Memory Network (Bi - LSTM) for feature extraction, and output a vector of the second dimension batch_size * max_seq_len * (2 * hidden_size).

[0106] The vector of the second dimension is then input into a conditional random field (CRF) model for decoding after passing through a fully connected layer. After passing through the fully connected layer, the size of the vector of the second dimension is transformed into batch_size*out_feature. The CRF model is a conditional probability model that can effectively handle the mutual constraint relationships between tags and effectively solve the sequence tagging problem. Therefore, the present application trains the model based on the scores of the tag sequences corresponding to the word-level sequences output by the CRF model and calculates the target annotation sequence.

[0107] To more clearly describe the training process of the BERT-Bi-LSTM-CRF model of the present application, each component in the BERT-Bi-LSTM-CRF model will be introduced first. As Figure 4 shown, the processing process of the Bert model is as follows: For any sequence, first obtain the text sequence in units of words, then mask some words in the sequence, and then add a special marker [CLS] at the beginning of the sequence, and separate sentences with the marker [SEP]. At this time, the output Embedding of each word in the sequence consists of three parts: TokenEmbedding, SegmentEmbedding, and PositionEmbedding. Then, the sequence vector is input into BERT for feature extraction, and finally a sequence vector with rich semantic features is obtained.

[0108] For BERT, its key part is the encoder part of the Transformer structure. The transformer uses a multi-head attention mechanism, which connects multiple attention layers with different initializations. The multi-head attention is expressed as follows:

[0109] Multihead(Q,K,V)=Concat(head 1 ,…,head n )W O

[0110] head i =Attention(QW i Q ,KW i K ,VW i KV )

[0111] where Denote the weight matrices corresponding to Q, K, and V. The Concat function represents concatenating the results of different heads. The Attention function can be described as mapping a query and a set of key-value pairs to an output, where the output is calculated as a weighted sum of the values, and the weight assigned to each value is calculated by a compatibility function of the query and the corresponding key. The basis of multi-head attention is the self-attention mechanism, and the principle is to adjust the weight coefficient matrix through the correlation degree between words in the same sentence to obtain the vector representation of words:

[0112]

[0113] Among them, Q, K, and V are vector matrices representing query, key, and value respectively, and d k is the vector dimension of the key.

[0114] In the embodiments of this application, the output of BERT is used as the input of the Bi-LSTM model to further extract context information. As Figure 5 shown, LSTM is a variant of RNN, which is an RNN with long short-term memory. Its network structure consists of an input gate, a forget gate, and an output gate, and it can perform better in longer sequences, overcoming the problems of gradient disappearance and gradient explosion existing in RNN. The formal representation of LSTM is as follows:

[0115] f t = σ(W f · [h t-1 , x t + b f )

[0116] i t = σ(W i · [h t-1 , x t + b i )

[0117]

[0118] o t = σ(W o · [h t-1 , x t + b o )

[0119] h t = o t ⊙ tanh(c t )

[0120] Among them, x t , h t-1 , ct-1 respectively represent the input at time t, the output at time t-1, and the cell state at t-1. ⊙ represents the dot product operation, σ is the sigmoid function, tanh represents the hyperbolic tangent function, and W * respectively represent the weights corresponding to the states, and b * represents the bias term.

[0121] In the embodiment of the present application, it is assumed that for the input sequence X (i.e., the word-level sequence), the out_feature obtained through the Bi-LSTM and the fully connected layer is an n*k matrix P, where n represents the length of the input sequence, and k represents the size of the label set. For the CRF, the input P i,j represents the probability that the i-th word in the input sequence corresponds to the j-th tag. By introducing the transition matrix T as a parameter of the CRF model, T i,j represents the transition probability from label i to label j for consecutive words. Then, for the input sequence X, the predicted label sequence y = {y 1 , y 2 , …, y n}, and the CRF model calculates the score of the label sequence y corresponding to the word-level sequence through the following formula:

[0122]

[0123] where X represents the word-level sequence, y represents the label sequence predicted by the CRF model for the word-level sequence, and P i,j represents the probability that the i-th word in the input word-level sequence corresponds to the j-th label in the label data set, and T i,j represents the transition probability from label i to label j for consecutive words in the word-level sequence.

[0124] Furthermore, after obtaining the score of the label sequence corresponding to the word-level sequence, the loss function is defined through the maximum log-likelihood function, that is, the loss function is defined as -log(p(y|X)), and the BERT-Bi-LSTM-CRF model is trained using the gradient descent method, where the maximum log-likelihood function is expressed as follows:

[0125]

[0126] where p(y|X) represents the probability value that each label sequence is the correct label sequence, and Y x represents all possible label sequences.

[0127] Even further, the target annotation sequence is calculated through the following formula:

[0128]

[0129] where y *It is the output sequence with the maximum conditional probability.

[0130] Thus, by training the model with the maximum likelihood function and adjusting the parameters of the model, after obtaining the output sequence with the maximum conditional probability from the maximum likelihood function, the model training is completed.

[0131] Step 106: Predict the results of multi-label sequence annotation through the trained BERT-Bi-LSTM-CRF model, and evaluate the prediction effect of the BERT-Bi-LSTM-CRF model.

[0132] In an embodiment of the present application, the results of sequence annotation are input into the trained BERT-Bi-LSTM-CRF model for prediction. The specific implementation manner of predicting the results of sequence annotation is as described in step 105, which will not be elaborated here, so as to achieve one-to-one correspondence for each segment of experience. Then, evaluate the prediction effect of the BERT-Bi-LSTM-CRF model.

[0133] When specifically evaluating the prediction effect of the model, as a possible implementation manner, first calculate the precision rate, recall rate, and comprehensive evaluation value of the answers generated by the fine-tuned BERT-Bi-LSTM-CRF model; then evaluate the fine-tuned BERT-Bi-LSTM-CRF model according to the precision rate, recall rate, and comprehensive evaluation value. Among them, the precision rate, recall rate, and comprehensive evaluation value are calculated through the following formulas:

[0134]

[0135] Among them,

[0136] Among them, P is the precision rate, R is the recall rate, F1 is the comprehensive evaluation value, m is the number of extracted records, n is the number of labeled records, and k is the number of elements of record i in the labeled data.

[0137] After calculating the precision rate, recall rate, and comprehensive evaluation value, the calculated values can be compared with a preset evaluation threshold. Among them, the evaluation threshold can be the lowest threshold when the information extraction effect of the preset model meets the requirements. By comparing, it is judged whether the above calculated values are greater than the evaluation threshold to evaluate whether the effect of the model meets the requirements.

[0138] In summary, the structured information extraction method based on the multi-labeling strategy in the embodiments of the present application first crawls the home pages of scholars, and performs cleaning processing and sentence splitting on the home pages; matches and converts dates in different formats in the sentences through regular expressions; filters out the texts containing the resume information of scholars through a preset short text classification model; performs multi-label sequence labeling on the texts based on the multi-labeling strategy to obtain a label data set, and splits the label data set into a training set, a validation set, and a test set. Train a Bidirectional Encoder Representations from Transformers-Bidirectional Long Short-Term Memory-Conditional Random Field (BERT-Bi-LSTM-CRF) model based on the data in the training set; predict the results of multi-label sequence labeling through the trained model. This method combines the multi-labeling strategy and the deep learning network model to perform structured extraction of the resume information of scholars, which can not only extract the basic information of scholars, but also solve the problem that the education or work experience of scholars does not match the timeline, and makes the information such as the institutions attended, majors, degrees, etc. of scholars correspond one by one with the start and end times to obtain accurate and structured data, meeting the application requirements of the actual scenario.

[0139] To more clearly illustrate the specific implementation process of the structured information extraction method based on the multi-labeling strategy, the following is combined with Figure 6 , and a specific embodiment is used for detailed description:

[0140] In this embodiment, in the first step, first crawl the home pages of experts and scholars, use the html2text toolkit to obtain the plain text pages, clean and split the sentences, and generate a data set containing personal profiles, including preprocessing work such as converting non-ASCII characters to Unicode characters, deleting illegal characters, converting traditional Chinese to simplified Chinese, and converting full-width and half-width characters, and then use the SentenceSpliter package of the Pyltp toolkit to split the text into sentences.

[0141] In the second step, use regular expressions to match various date formats in the text obtained by splitting sentences, summarize various date formats to formulate a regular expression with a high date coverage rate, extract information such as the year and month, and then convert it into a unified time format of YYYY year or YYYY year-MM month.

[0142] In the third step, a short text classification model is constructed based on TextCNN (including a vector representation layer, a convolutional layer, a max pooling layer, and a fully connected layer) to filter out texts containing scholars' resume information. For example, texts containing scholars' educational or work experiences are filtered out. In the vector representation layer, each character in the filtered text is mapped into a 100-dimensional vector through word2vec. Then, the convolutional layer is used to extract information in the sentence with different kernel sizes. Next, the max pooling layer takes the maximum value of several one-dimensional vectors obtained after convolution and then concatenates them together as the output value of this layer. Finally, in the fully connected layer, softmax is used for classification. The input is the pooled vector, and the output dimension is 3, representing the three categories of education, work, and others. And to prevent overfitting, an L2 regularization and dropout are introduced.

[0143] In the fourth step, a multi-label annotation strategy is adopted to annotate the data in the text to generate a dataset. That is, a multi-label layer is set, and all labels are given to each token according to each experience. Then these label layers are merged, and finally, they are split according to the ratio of 7:1.5:1.5 to generate a training set, a validation set, and a test set.

[0144] In the fifth step, a BERT-Bi-LSTM-CRF structured information extraction model is trained. Specifically, in implementation, the BERT-Bi-LSTM-CRF model can be used for sequence annotation. The input is the tokenid obtained by wordPiecetokenizer, which enters the Bert pre-trained model to extract rich text features to obtain an output vector of batch_size*max_seq_len*emb_size. The Bi-LSTM is used to extract the features required for entity recognition to obtain a vector of batch_size*max_seq_len*(2*hidden_size). After passing through the fully connected layer, it finally enters the CRF layer for decoding to calculate the optimal annotation sequence.

[0145] In the sixth step, the results of multi-label sequence annotation are predicted, the data format is converted into structured data, so that each experience corresponds one by one, and the results of each experience prediction are evaluated.

[0146] To implement the above embodiments, the present application also proposes a structured information extraction device based on a multi-label annotation strategy.

[0147] Figure 7 It is a schematic structural diagram of a structured information extraction device based on a multi-label annotation strategy proposed for the embodiments of the present application.

[0148] Such as Figure 7As shown in the figure, the structured information extraction device based on the multi-labeling strategy includes a homepage generation module 100, a date generation module 200, a sentence classification module 300, a multi-labeling module 400, a model training module 500, and a prediction module 600.

[0149] Among them, the homepage generation module 100 is used to obtain the homepage of the scholar and perform cleaning processing and sentence splitting processing on the homepage.

[0150] The date generation module 200 is used to match each sentence through regular expressions to obtain dates in different formats on the homepage and convert the format of each date into a preset date format.

[0151] The sentence classification module 300 is used to classify each sentence through a preset short text classification model and screen out the text representing the educational or work experience of the scholar.

[0152] The multi-labeling module 400 is used to perform multi-label sequence labeling on the text based on the multi-labeling strategy to obtain a label data set, and divide the label data set into a training set, a validation set, and a test set according to a preset ratio.

[0153] The model training module 500 is used to train a Bidirectional Encoder Representations from Transformers - Bidirectional Long Short-Term Memory Network - Conditional Random Field (BERT - Bi - LSTM - CRF) model based on the data in the training set.

[0154] The prediction module 600 is used to predict the results of multi-label sequence labeling through the trained BERT - Bi - LSTM - CRF model and evaluate the prediction effect of the BERT - Bi - LSTM - CRF model.

[0155] Optionally, in an embodiment of the present application, the homepage generation module 100 is specifically used to: convert the page language of the homepage into a lightweight markup language Markdown; convert non - ASCII characters in the homepage into Unicode characters and delete illegal characters; convert the font and case of each character in the homepage; perform sentence splitting processing on the homepage, including: dividing the content in the homepage into multiple sentences through a preset toolkit.

[0156] Optionally, in an embodiment of the present application, the sentence classification module 300 is specifically used to: map each character in each sentence into a vector of a preset dimension; perform convolution processing on each sentence through different convolution kernel sizes to obtain multiple one - dimensional column vectors; take the maximum value in each one - dimensional column vector and splice each maximum value to obtain a pooled vector; classify the pooled vector through the softmax function to determine the category to which each sentence belongs.

[0157] Optionally, in an embodiment of the present application, the multi-annotation module 400 is specifically configured to: perform position part annotation on each word in the text through BIO sequence annotation; perform entity type part annotation on each word in the text; set a number for each piece of experience, and perform experience stage part annotation on each word in the text.

[0158] Optionally, in an embodiment of the present application, the model training module 500 is specifically configured to: convert each sentence in the training set into a word-level sequence, and input each word-level sequence into the BERT model to generate a vector of the first size based on context information; input the vector of the first size into a bidirectional long short-term memory network Bi-LSTM for feature extraction, and output a vector of the second size; input the vector of the second size through a fully connected layer into a conditional random field CRF model for decoding, train the model based on the scores of the label sequences corresponding to the word-level sequences output by the CRF model, and calculate the target annotation sequence.

[0159] It should be noted that, in an embodiment of the present application, the CRF model calculates the scores of the label sequences corresponding to the word-level sequences through the following formula:

[0160]

[0161] where X represents the word-level sequence, y represents the label sequence predicted by the CRF model for the word-level sequence, P i,j represents the probability that the i-th word in the input word-level sequence corresponds to the j-th label in the label dataset, and T i,j represents the transition probability from label i to label j for consecutive words in the word-level sequence.

[0162] It should be noted that the model training module 500 is specifically configured to, after obtaining the scores of the label sequences corresponding to the word-level sequences, define a loss function through the maximum log-likelihood function, and use the gradient descent method to train the BERT-Bi-LSTM-CRF model. The maximum log-likelihood function is expressed as follows:

[0163]

[0164] where p(y|X) represents the probability value that each label sequence is the correct label sequence, and Y x represents all possible label sequences;

[0165] The target annotation sequence is calculated through the following formula:

[0166]

[0167] where y * is the output sequence with the maximum conditional probability.

[0168] Optionally, in an embodiment of the present application, the prediction module 600 is specifically configured to: calculate the precision rate, recall rate, and comprehensive evaluation value of the answers generated by the fine-tuned training model; evaluate the fine-tuned generative training model according to the precision rate, recall rate, and comprehensive evaluation value; wherein, the precision rate, recall rate, and comprehensive evaluation value are calculated by the following formulas:

[0169]

[0170] Wherein,

[0171] Wherein, P is the precision rate, R is the recall rate, F1 is the comprehensive evaluation value, m is the number of extracted records, n is the number of labeled records, and k is the number of elements of record i in the labeled data.

[0172] In summary, the structured information extraction device based on the multi-labeling strategy in the embodiment of the present application first crawls the homepage of the scholar, and performs cleaning processing and sentence splitting processing on the homepage; matches and converts dates in different formats in the sentence through regular expressions; filters out texts representing the education or work experience of the scholar through a preset short text classification model; performs multi-label sequence labeling on the text based on the multi-labeling strategy to obtain a label data set, and divides the label data set into a training set, a validation set, and a test set. Train a Bidirectional Encoder Representations from Transformers-Bidirectional Long Short-Term Memory-Conditional Random Field (BERT-Bi-LSTM-CRF) model based on the data in the training set; predict the results of multi-label sequence labeling through the trained model. The device combines the multi-labeling strategy and the deep learning network model to perform structured extraction of the scholar's resume information, which can not only extract the basic information of the scholar, but also solve the problem of the mismatch between the education or work experience of the scholar and the timeline, and correspond the information such as the scholar's school, major, degree, etc. with the start and end times one by one to obtain accurate and structured data, meeting the application requirements of the actual scenario.

[0173] To implement the above embodiment, the present invention also proposes a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the structured information extraction method based on the multi-labeling strategy described in the above embodiment of the present application.

[0174] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, if the schematic expressions of the above terms are used in multiple embodiments or examples, it does not mean that these embodiments or examples are the same. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0175] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0176] Any process or method description shown in a flowchart or described in other ways herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0177] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0178] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0179] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0180] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0181] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.

Claims

1. A structured information extraction method based on a multi - annotation strategy, characterized in that, it includes the following steps: Crawl the homepage of the scholar, and perform cleaning processing and sentence splitting processing on the homepage; Match each sentence through regular expressions to obtain dates in different formats on the homepage, and convert the format of each date into a preset date format; Classify each sentence through a preset short - text classification model, and screen out the text containing the scholar's resume; Perform multi - label sequence annotation on the text based on the multi - annotation strategy to obtain a label data set, and divide the label data set into a training set, a validation set and a test set according to a preset ratio; Train a Bidirectional Encoder Representations from Transformers - Bidirectional Long Short - Term Memory Network - Conditional Random Field (BERT - Bi - LSTM - CRF) model based on the data in the training set; Predict the results of multi - label sequence annotation through the trained BERT - Bi - LSTM - CRF model, and evaluate the prediction effect of the BERT - Bi - LSTM - CRF model; Among them, the evaluation of the prediction effect of the BERT - Bi - LSTM - CRF model includes: Calculate the precision rate, recall rate and comprehensive evaluation value of the results predicted by the trained BERT - Bi - LSTM - CRF model; Evaluate the trained BERT - Bi - LSTM - CRF model according to the precision rate, the recall rate and the comprehensive evaluation value; among them, the precision rate, the recall rate and the comprehensive evaluation value are calculated through the following formulas: Among them, Where P is the precision rate, R is the recall rate, F1 is the comprehensive evaluation value, m is the number of extracted records, n is the number of annotated records, and k is the number of elements of record i in the annotated data.

2. The extraction method according to claim 1, characterized in that, The cleaning process of the homepage includes: Convert the page language of the homepage into the lightweight markup language Markdown; Convert non - ASCII characters in the homepage into Unicode characters and delete illegal characters; Convert the font and case of each character in the homepage; The sentence splitting process of the homepage includes: Divide the content in the homepage into multiple sentences through a preset toolkit.

3. The extraction method according to claim 1, characterized in that, The short - text classification model includes a vector representation layer, a convolutional layer, a max - pooling layer and a fully - connected layer. The classification of each sentence through the preset short - text classification model includes: Map each character in each sentence into a vector of a preset dimension; Perform convolution processing on each sentence through different convolutional kernel sizes to obtain multiple one - dimensional column vectors; Take the maximum value in each one - dimensional column vector and splice each maximum value to obtain a pooled vector; Classify the pooled vector through the softmax function to determine the category to which each sentence belongs.

4. The extraction method according to claim 1, characterized in that, Performing sequence labeling on the text based on the multi - label annotation strategy includes: Performing position - part labeling on each character in the text through BIO sequence labeling; Performing entity - type - part labeling on each character in the text; Setting a number for each piece of experience and performing experience - stage - part labeling on each character in the text.

5. The extraction method according to claim 1, wherein, Training a Transformer - based Bidirectional Encoder Representations from Transformers - Bidirectional Long Short - Term Memory Network - Conditional Random Field (BERT - Bi - LSTM - CRF) model based on the data in the training set includes: Converting each sentence in the training set into a character - level sequence and inputting each character - level sequence into the BERT model to generate a vector of the first size based on context information; Inputting the vector of the first size into a Bidirectional Long Short - Term Memory Network (Bi - LSTM) for feature extraction and outputting a vector of the second size; Inputting the vector of the second size through a fully - connected layer into a Conditional Random Field (CRF) model for decoding, training the model based on the scores of the label sequence corresponding to the character - level sequence output by the CRF model, and calculating the target annotation sequence.

6. The extraction method according to claim 5, wherein, The CRF model calculates the scores of the label sequence corresponding to the character - level sequence through the following formula: where X represents the word sequence, y represents the label sequence after the CRF model predicts the word-level sequence, and P i,j represents the probability that the i-th word in the input word-level sequence corresponds to the j-th label in the label dataset, and T i,j represents the transition probability of consecutive words in the word-level sequence from label i to label j.

7. The extraction method according to claim 6, wherein, After obtaining the scores of the label sequence corresponding to the character - level sequence, defining a loss function through the maximum log - likelihood function and training the BERT - Bi - LSTM - CRF model using the gradient descent method. The maximum log - likelihood function is expressed as follows: Among them, p(y|X) represents the probability value that each label sequence is the correct label sequence, and Y x represents all possible label sequences; Calculating the target annotation sequence through the following formula: where y * is the output sequence with the maximum conditional probability.

8. A structured information extraction device based on a multi - label annotation strategy, wherein, comprises: A homepage generation module for obtaining the homepage of a scholar and performing cleaning processing and sentence - splitting processing on the homepage; A date generation module for matching each sentence through a regular expression to obtain dates in different formats in the homepage and converting the format of each date into a preset date format; A sentence classification module for classifying each sentence through a preset short - text classification model and screening out the text containing the scholar's resume; A multi - label annotation module for performing multi - label sequence labeling on the text based on a multi - label annotation strategy to obtain a label data set, and splitting the label data set into a training set, a validation set, and a test set according to a preset ratio; A model training module for training a Transformer - based Bidirectional Encoder Representations from Transformers - Bidirectional Long Short - Term Memory Network - Conditional Random Field (BERT - Bi - LSTM - CRF) model based on the data in the training set; A prediction module for predicting the results of multi - label sequence labeling through the trained BERT - Bi - LSTM - CRF model and evaluating the prediction effect of the BERT - Bi - LSTM - CRF model; Among them, the prediction module is specifically configured to calculate the precision rate, recall rate, and comprehensive evaluation value of the result predicted by the trained BERT-Bi-LSTM-CRF model; Evaluate the trained BERT-Bi-LSTM-CRF model according to the precision rate, the recall rate, and the comprehensive evaluation value; among them, the precision rate, the recall rate, and the comprehensive evaluation value are calculated by the following formulas: Among them, Among them, P is the precision rate, R is the recall rate, F1 is the comprehensive evaluation value, m is the number of extracted records, n is the number of labeled records, and k is the number of elements of record i in the labeled data.

9. A non-transitory computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it implements the structured information extraction method based on a multi-labeling strategy according to any one of claims 1-7.

Citation Information

Patent Citations

  • Neural network-based scholarship user portrait information extraction method and model

    CN109657135A