Entity extraction model training method, network equipment and storage medium
By using large models for data screening and incremental training in the entity extraction model in the communication field, the problems of few entity training samples, difficult labeling, high training cost and low entity extraction accuracy are solved, and a more efficient training process and more accurate entity extraction effect are achieved.
Patent Information
- Application Number
- CN202311455994.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-13
AI Technical Summary
The entity extraction model in the communication field has problems such as few entity training samples, difficulty in labeling, high training costs and low entity extraction accuracy.
By obtaining the first training data and the first entity extraction model for training, the second entity extraction model is obtained; then obtaining the second training data for information extraction, and the first prediction data is obtained; based on the data screening model, the third training data is obtained; finally, the second entity extraction model is incrementally trained to obtain the target entity extraction model.
Use large models to quickly filter data, obtain enhanced sample data, and through data augmentation and model iterative training, we train the entity extraction model that meets the needs, reduce training costs, and improve the accuracy of entity extraction.
Smart Images

Figure CN119988958A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a training method, network equipment and storage medium for an entity extraction model. Background Art
[0002] With the development of artificial intelligence technology, specific information extraction work on text knowledge in the field of communications has received more and more attention. Although small sample models can achieve satisfactory results in general fields based on a small number of labeled samples; however, they cannot achieve satisfactory results in the field of communications. Due to the large number of proper nouns in the field of communications, it is difficult for pre-trained models to identify proper nouns well. Models trained using only small sample annotations often cannot cover all proper noun types and expression paradigms, and the model recall rate often cannot meet actual business needs. Therefore, the entity extraction model in the relevant technology has problems such as few proper noun entity training samples, difficult labeling, high training costs, and low entity extraction accuracy. Summary of the invention
[0003] The main purpose of the embodiments of the present application is to provide a training method, network device and storage medium for an entity extraction model, aiming to solve the problems of small number of entity training samples, difficult labeling, high training cost and low entity extraction accuracy in the entity extraction model in the communication field.
[0004] In a first aspect, an embodiment of the present application provides a method for training an entity extraction model, the method comprising:
[0005] Acquire first training data and a first entity extraction model, and train the first entity extraction model according to the first training data to obtain a second entity extraction model;
[0006] Acquire second training data, and extract information from the second training data using the second entity extraction model to obtain first prediction data;
[0007] Determine entity extraction template information and entity extraction restriction information according to the second training data, and based on a data screening model, screen the first prediction data according to the entity extraction template information and the entity extraction restriction information to obtain third training data;
[0008] Incrementally train the second entity extraction model according to the third training data to obtain a target entity extraction model.
[0009] In a second aspect, an embodiment of the present application provides a network device, the network device comprising:
[0010] A processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing connection and communication between the processor and the memory, wherein when the computer program is executed by the processor, the steps of the training method of the entity extraction model as described above are implemented.
[0011] In a third aspect, an embodiment of the present application provides a storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the training method of the entity extraction model as described above.
[0012] The present application provides a training method, network device and storage medium for an entity extraction model, which obtains first training data and a first entity extraction model, and trains the first entity extraction model according to the first training data to obtain a second entity extraction model; obtains second training data, and extracts information from the second training data through the second entity extraction model to obtain first prediction data; determines entity extraction template information and entity extraction restriction information according to the second training data, and based on a data screening model, screens the first prediction data according to the entity extraction template information and entity extraction restriction information to obtain third training data; performs incremental training on the second entity extraction model according to the third training data to obtain a target entity extraction model. In this way, a large model can be used to realize rapid screening of data to obtain enhanced sample data, so that the characteristics of the proper noun data in the field of communication operation and maintenance can be fully utilized, and an entity extraction model that meets the requirements can be trained by data enhancement and model iterative training in the case of using a small number of samples. The process of training the target entity extraction model is simpler and more efficient, and the training cost is reduced, and the entity extraction accuracy of the target entity extraction model is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1 A flowchart of a method for training an entity extraction model provided in an embodiment of the present application;
[0015] Figure 2 A flowchart of the overall process of a training method for an entity extraction model provided in an embodiment of the present application;
[0016] Figure 3A schematic block diagram of the structure of a training device for an entity extraction model provided in an embodiment of the present application;
[0017] Figure 4 A schematic block diagram of the structure of a network device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0019] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0020] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0021] Due to the professionalism of the communications field, labeled data in the communications field is very scarce. At the same time, manually labeling training data is very expensive. In the process of building a knowledge graph, it is inevitable to train the model of the entity information involved in the schema. At the same time, in communications operation and maintenance, the domain's proper nouns often play a very important role in the process of building a knowledge graph. The extraction of domain proper nouns can greatly improve the use value of the knowledge graph to a certain extent.
[0022] However, the rapid development of natural language processing technology, especially the development of large-scale pre-trained language models, has played a vital role in the information extraction technology in knowledge graphs. However, pre-trained models usually use corpus data from large-scale public data sets, so it is difficult to directly apply them to vertical fields such as communication operations.
[0023] The development of small-shot model training technology enables the model to achieve satisfactory results in general fields based on a small number of labeled samples. Therefore, the current mainstream information extraction model for proper nouns usually uses small-shot technology, that is, fine-tuning the pre-trained model (such as BERT, RoBERTa, ERNIE and other encoder models) by manually annotating a small number of samples (several or dozens).
[0024] However, there are many proper nouns in the field of communications, and it is difficult for pre-trained models to identify proper nouns well. Models trained simply using small sample annotations often cannot cover all types of proper nouns and expression paradigms, and the model recall rate often cannot meet actual business needs.
[0025] The embodiment of the present application provides a training method, network device and storage medium for an entity extraction model. The method can be applied to network devices, wherein the network devices can be servers, base stations, switches, routers, telecommunications network management products and knowledge graph products, etc., but are not limited thereto. Thus, a large model can be used to quickly screen data to obtain enhanced sample data, so that the characteristics of proper noun data in the field of communication operation and maintenance can be fully utilized. In the case of using a small number of samples, an entity extraction model that meets the requirements can be trained through data enhancement and model iterative training. The process of training the target entity extraction model is simpler and more efficient, and the training cost is reduced, and the entity extraction accuracy of the target entity extraction model is improved.
[0026] In summary, the entity extraction model training methods in related technologies have at least the following problems: (1) The training data in the communication field is limited, and the model obtained by training only with a small sample has low accuracy; (2) The data structure constructed by the traditional rule-based matching method is single, and the manual definition of rules is cumbersome; (3) The semantic-based data expansion method is often subject to the accuracy of the similarity matching algorithm, resulting in the problem of incoherent semantics of the expanded samples.
[0027] Please refer to Figure 1 , Figure 1 A flow chart of a training method for an entity extraction model provided in an embodiment of the present application. The training method for the entity extraction model can use a large model to achieve rapid screening of data to obtain enhanced sample data, thereby being able to fully utilize the characteristics of proper noun data in the field of communication operation and maintenance, and train an entity extraction model that meets the requirements through data enhancement and model iterative training in the case of using a small number of samples, making the process of training the target entity extraction model simpler and more efficient, reducing training costs, and improving the entity extraction accuracy of the target entity extraction model.
[0028] like Figure 1As shown, the training method of the entity extraction model may include S101 to S104.
[0029] S101, obtaining first training data and a first entity extraction model, and training the first entity extraction model according to the first training data to obtain a second entity extraction model.
[0030] The first training data includes text data after entity annotation, and the text data after entity annotation may include entity information, and the entity information may be a term specific to the communication field, such as parameters, alarm indicators, and performance indicators, etc. The first entity extraction model is an untrained initial entity extraction model, and the entity extraction model may be a small sample model, such as encoder models such as BERT, RoBERTa, and ERNIE. The small sample model can achieve satisfactory results in general fields based on a small number of annotated samples. The second entity extraction model is the trained first entity extraction model.
[0031] In some embodiments, text data is obtained and segmented to obtain multiple text information segments; the annotation format corresponding to each text information segment is determined, and the corresponding text information is annotated according to the annotation format corresponding to each text information segment to obtain the first training data. Thus, the first training data can be generated in advance to train the first entity extraction model to obtain the second entity extraction model.
[0032] The text data may be text data to be extracted in the communication field. The text data to be extracted is usually in the form of long text, mainly in the form of fault cases, document manuals, expert knowledge, web page data, experience summaries, etc. The text information is the text obtained by segmentation from the text data. The annotation format may include BIO annotation format or pointer annotation format, etc.
[0033] Exemplarily, the marked entity information can be a term in the field of communications, such as parameters, alarm indicators, and performance indicators, etc. Parameters can include mechanical downtilt angle, center frequency, PRACH parameters, and A2 threshold, etc.; alarm indicators can include S1 link break alarm, performance threshold crossing alarm, standing wave ratio alarm, and RRU internal fault alarm, etc.; performance indicators can include the number of RRC connection establishment requests, same-frequency switching rate, cell wireless connection rate, and number of dropped calls, etc.
[0034] For example, if the text data is long text data, it can be expressed in the following form:
[0035] [Fault phenomenon] The uplink interference of the three cells in Machangyaoshangtun_ZLH has increased as a whole, and the interference level is about -82dBm. The interference curve analysis of RB0-RB99 shows a full-frequency increase, as shown in the figure below. There are also uplink interference in the other four cells around. [Cause analysis] 1. The seven cells have full-band interference increase, and the interference intensity of each RB is equivalent, which can rule out the problem of interference from different systems and spurious interference; 2. The alarm check also did not find GPS lock-out alarms; 3. The cell with the strongest interference is turned off, and interference from other cells still exists. Therefore, the possibility of being interfered by external jammers is the greatest, and it is necessary to scan the frequency on site for investigation. [Problem solution] On the afternoon of January 20, I went to the Chengzhong Village of Yaoshangtun with the Wireless Committee to scan the frequency on site, and located the interference source as a privately installed signal blocker.
[0036] The long text data may be segmented by symbols or paragraphs to obtain a plurality of short text data, and part of the short text data may be selected for entity annotation to obtain the first training data.
[0037] Exemplarily, the annotation format can be determined according to the specific model of the first entity extraction model used. For example, if the first entity extraction model is a BEET+CRF model, the annotation format can adopt the BIO annotation format; if the first entity extraction model is a PaddleNLP model, the annotation format can adopt the pointer annotation format.
[0038] S102: Acquire second training data, and extract information from the second training data using a second entity extraction model to obtain first prediction data.
[0039] The second training data includes text data without entity annotation. The first prediction data is entity information obtained after the second training data is subjected to information extraction by the second entity extraction model.
[0040] In some embodiments, the second training data is subjected to information extraction by a second entity extraction model to obtain entity recognition information corresponding to the second training data; the confidence of each entity recognition information is determined, and each entity recognition information is screened according to the confidence of each entity recognition information to obtain the first prediction data. In this way, information extraction can be performed accurately and the first prediction data can be obtained by screening.
[0041] The entity recognition information may be entity information obtained after the second training data is subjected to entity extraction by a second entity extraction model.
[0042] Specifically, information is extracted from the second training data through the second entity extraction model to obtain entity recognition information corresponding to multiple segments of text information in the second training data; the confidence of each entity recognition information is determined, and each entity recognition information is screened according to the confidence of each entity recognition information, and finally the entity recognition information that meets the confidence requirement is integrated to obtain the first prediction data.
[0043] In some embodiments, it is determined whether the confidence of each entity identification information exceeds a preset confidence; if the confidence of the entity identification information exceeds the preset confidence, the entity identification information is used as the first prediction data, thereby accurately extracting information and screening to obtain the first prediction data.
[0044] The preset confidence level may be a preset confidence level standard value, such as 0.9, and may be set according to actual conditions, and is not specifically limited here.
[0045] Specifically, determine whether the confidence of each entity recognition information exceeds the preset confidence; if the confidence of the entity recognition information exceeds the preset confidence, use the entity recognition information as the first prediction data; if the confidence of the entity recognition information does not exceed the preset confidence, filter out the entity recognition information.
[0046] Exemplarily, if the preset confidence is 0.9, and the entity recognition information corresponding to the multiple text information in the second training data includes entity recognition information A, entity recognition information B, and entity recognition information C, the confidences of entity recognition information A, entity recognition information B, and entity recognition information C are determined to be 0.95, 0.92, and 0.8, respectively. At this time, it can be determined that the confidences of entity recognition information A and entity recognition information B exceed the preset confidence, and entity recognition information C does not exceed the preset confidence, so entity recognition information A and entity recognition information B can be used as the first prediction data.
[0047] S103: Determine entity extraction template information and entity extraction restriction information according to the second training data, and based on the data screening model, screen the first prediction data according to the entity extraction template information and the entity extraction restriction information to obtain third training data.
[0048] Among them, the entity extraction template information can be sample template information related to proper nouns in the communication field. The entity extraction restriction information can be input and output restriction conditions for entity extraction. The data screening model can be a large model, such as ChatGPT, etc., which is used to perform NLP tasks and perform data screening. The third training data can be the first prediction data obtained by filtering the data screening model.
[0049] In some embodiments, the task type and task domain corresponding to the second training data are determined, and the entity template information and entity restriction information are determined according to the task type and task domain corresponding to the second training data, thereby accurately determining the entity template information and entity restriction information.
[0050] The task type corresponding to the second training data can be used to represent the extraction object type of the entity extraction model, such as a proper noun type, a common noun type, or a name or place name type, etc. The task field corresponding to the second training data can be used to represent the extraction object field of the entity extraction model, such as a telecommunications network management knowledge graph field, a telecommunications network management entity extraction field, or a natural language analysis field, etc.
[0051] Specifically, semantic recognition is performed on the content of the text data to obtain semantic recognition information corresponding to the text data; the task type and task domain corresponding to the second training data are determined based on the semantic recognition information; and entity template information and entity restriction information are determined based on the task type and task domain corresponding to the second training data.
[0052] Among them, the entity extraction template information may include instruction description information, proper noun explanation information and sample example information. The instruction description information is used to define the task type, the proper noun explanation information is used to provide the explanation of the part of speech, data type, noun meaning, functional module, etc. of the proper nouns to be extracted by the telecommunications network management, and the sample example information is used to provide the extraction results of a small amount of text to be extracted, and at the same time provide the result judgment description after the information extraction. The entity extraction restriction information may include role definition information, input filling information, output limitation information and output instruction information. The role definition information is used to define the data screening model in the corresponding field proper noun entity extraction role, the input filling information is used to complete the replacement and filling of the input data, and the output limitation information is used to limit the output content of the data screening model, and the limitation includes but is not limited to the output format, the length of the output text, the output content constraints, etc.; the output instruction information is used to instruct the large model to perform output parsing.
[0053] Exemplarily, if the task type corresponding to the second training data is a proper noun type, and the task domain is a telecommunications network management entity extraction domain, the role definition information may be "a telecommunications network management entity extraction expert"; the instruction description information may be "determine whether the content enclosed in Chinese brackets in the input text is a proper noun type entity"; the proper noun explanation information may provide different proper noun explanation files in the telecommunications network management entity extraction domain to be filled into the corresponding prompt word template; the sample example information may provide sample information corresponding to the proper noun extraction.
[0054] For example, the entity extraction template information and the entity extraction restriction information can be integrated into a table to represent the entity extraction template information and the entity extraction restriction information.
[0055] As shown in Table 1.
[0056]
[0057]
[0058] Table 1
[0059] It can be seen that Table 1 can represent entity extraction template information and entity extraction restriction information, so that the data screening model can screen the first prediction data based on various entity extraction template information and entity extraction restriction information in Table 1 to obtain the third training data.
[0060] In some embodiments, a data output template corresponding to the data screening model is determined according to the entity template information; the first prediction data is screened according to the entity extraction restriction information, and the screened first prediction data is integrated based on the data output template to obtain the third training data. In this way, the third training data that meets the corresponding template can be accurately screened and output.
[0061] The data output template may include the data output format and data output style of the data screening model, etc.
[0062] Specifically, the data output format and data output style corresponding to the data screening model can be determined based on information such as instruction description information, term explanation information, and sample example information. The first predicted data is first preliminarily screened according to the entity extraction restriction information to obtain the first predicted data that meets the entity extraction restriction information. Finally, based on the data output template, the filtered first predicted data is integrated and processed to obtain the third training data.
[0063] S104: incrementally train the second entity extraction model according to the third training data to obtain a target entity extraction model, wherein the target entity extraction model is the second entity extraction model after incremental training using the third training data, and is an entity extraction model that can meet the user's entity extraction requirements.
[0064] In some embodiments, the second entity extraction model is incrementally trained according to the third training data to obtain a third entity extraction model; the entity extraction accuracy of the third entity extraction model is determined; if the entity extraction accuracy of the third entity extraction model exceeds a preset accuracy threshold, the third entity extraction model is used as the target entity extraction model. In this way, an entity extraction model that meets the user's entity extraction needs can be generated.
[0065] The third entity extraction model is a second entity extraction model that is incrementally trained using the third training data. The entity extraction accuracy rate is used to represent the accuracy rate of entity information obtained by entity extraction by the third entity extraction model. The preset accuracy rate threshold may be a preset entity extraction accuracy rate, such as 80%, which may be determined based on actual conditions.
[0066] Specifically, the second entity extraction model is incrementally trained according to the third training data to obtain a third entity extraction model; test data is obtained, and the entity extraction accuracy of the third entity extraction model is determined according to the test data; if the entity extraction accuracy of the third entity extraction model exceeds a preset accuracy threshold, the third entity extraction model is used as the target entity extraction model; if the entity extraction accuracy of the third entity extraction model does not exceed the preset accuracy threshold, new first training data can be re-acquired, and the first entity extraction model is trained with the new first training data to obtain the second entity extraction model, that is, step S101 is re-executed.
[0067] Exemplarily, if the preset accuracy threshold is 90%, if the entity extraction accuracy of the third entity extraction model is determined to be 95%, it can be determined that the entity extraction accuracy of the third entity extraction model exceeds the preset accuracy threshold, and the third entity extraction model is used as the target entity extraction model.
[0068] In some embodiments, after determining the entity extraction accuracy of the third entity extraction model, if the entity extraction accuracy of the third entity extraction model does not exceed the preset accuracy threshold, the second training data is extracted through the third entity extraction model to obtain second prediction data; based on the data screening model, the second prediction data is screened according to the entity extraction template information and the entity extraction restriction information to obtain fourth training data; the third entity extraction model is incrementally trained according to the fourth training data until the entity extraction accuracy of the third entity extraction model after incremental training exceeds the preset accuracy threshold. In this way, iterative model training can be performed to obtain an entity extraction model that can meet the user's entity extraction needs.
[0069] The second prediction data is entity information obtained after the second training data is subjected to information extraction by the third entity extraction model. The fourth training data can be training data obtained after the second prediction data is screened by the data screening model.
[0070] Specifically, if the entity extraction accuracy of the third entity extraction model does not exceed the preset accuracy threshold, the second training data is extracted through the third entity extraction model to obtain the second prediction data again; based on the data screening model, the second prediction data is screened again according to the entity extraction template information and the entity extraction restriction information to obtain the fourth training data; the third entity extraction model is incrementally trained according to the fourth training data, and it is determined whether the entity extraction accuracy of the third entity extraction model after the incremental training exceeds the preset accuracy threshold; if the entity extraction accuracy of the third entity extraction model after the incremental training exceeds the preset accuracy threshold, the iterative training is stopped, and the third entity extraction model after the incremental training is used as the target entity extraction model; if the entity extraction accuracy of the third entity extraction model after the incremental training does not exceed the preset accuracy threshold, the above training process is repeated until the entity extraction accuracy of the third entity extraction model after multiple incremental trainings exceeds the preset accuracy threshold.
[0071] In some embodiments, Figure 2 As shown, the overall process of entity extraction model training is introduced to introduce the training method of the entity extraction model proposed in one embodiment of the present application, including S201-S207.
[0072] S201. Train a first entity extraction model according to first training data to obtain a second entity extraction model.
[0073] S202: extract information from the second training data using a second entity extraction model to obtain first prediction data.
[0074] S203: Based on the data screening model, the first prediction data is screened according to the entity extraction template information and the entity extraction restriction information to obtain third training data.
[0075] S204: Incrementally train the second entity extraction model according to the third training data to obtain a third entity extraction model.
[0076] S205, determining whether the entity extraction accuracy of the third entity extraction model exceeds a preset accuracy threshold; if so, proceed to S207; if not, proceed to S206;
[0077] S206: Replace the second entity extraction model with the third entity extraction model.
[0078] S207: Use the third entity extraction model after incremental training as the target entity extraction model.
[0079] The present application provides a training method for an entity extraction model, which obtains a second entity extraction model by obtaining first training data and a first entity extraction model, and training the first entity extraction model according to the first training data; obtaining second training data, and extracting information from the second training data through the second entity extraction model to obtain first prediction data; determining entity extraction template information and entity extraction restriction information according to the second training data, and based on a data screening model, screening the first prediction data according to the entity extraction template information and entity extraction restriction information to obtain third training data; performing incremental training on the second entity extraction model according to the third training data to obtain a target entity extraction model. In this way, a large model can be used to realize rapid screening of data to obtain enhanced sample data, so that the characteristics of the proper noun data in the field of communication operation and maintenance can be fully utilized, and an entity extraction model that meets the requirements can be trained by data enhancement and model iterative training in the case of using a small number of samples. The process of training the target entity extraction model is simpler and more efficient, and the training cost is reduced, and the entity extraction accuracy of the target entity extraction model is improved.
[0080] See also Figure 3 , Figure 3 It is a schematic block diagram of a training device for an entity extraction model provided in one embodiment of the present application. The training device for the entity extraction model can be configured in a server to execute the aforementioned training method for the entity extraction model.
[0081] like Figure 3 As shown, the training device 300 of the entity extraction model includes: a first model training module 301, an entity extraction module 302, a data screening module 303 and a second model training module 304.
[0082] A first model training module 301 is used to obtain first training data and a first entity extraction model, and train the first entity extraction model according to the first training data to obtain a second entity extraction model;
[0083] An entity extraction module 302 is used to obtain second training data and extract information from the second training data using the second entity extraction model to obtain first prediction data;
[0084] A data screening module 303, configured to determine entity extraction template information and entity extraction restriction information according to the second training data, and based on a data screening model, screen the first prediction data according to the entity extraction template information and the entity extraction restriction information to obtain third training data;
[0085] The second model training module 304 is used to perform incremental training on the second entity extraction model according to the third training data to obtain a target entity extraction model.
[0086] The data screening module 303 is further used to determine the task type and task domain corresponding to the second training data; and determine the entity template information and entity restriction information according to the task type and task domain corresponding to the second training data.
[0087] The data screening module 303 is also used to determine the data output template corresponding to the data screening model according to the entity template information; screen the first prediction data according to the entity extraction restriction information, and integrate the screened first prediction data based on the data output template to obtain third training data.
[0088] The first model training module 301 is also used to obtain text data and segment the text data to obtain multiple segments of text information; determine the annotation format corresponding to each segment of the text information, and annotate the corresponding text information according to the annotation format corresponding to each segment of the text information to obtain the first training data.
[0089] The entity extraction module 302 is also used to extract information from the second training data through the second entity extraction model to obtain entity recognition information corresponding to the second training data; determine the confidence of each entity recognition information, and filter each entity recognition information according to the confidence of each entity recognition information to obtain first prediction data.
[0090] The entity extraction module 302 is further used to determine whether the confidence of each entity recognition information exceeds a preset confidence; if the confidence of the entity recognition information exceeds the preset confidence, the entity recognition information is used as the first prediction data.
[0091] The second model training module 304 is also used to perform incremental training on the second entity extraction model according to the third training data to obtain a third entity extraction model; determine the entity extraction accuracy of the third entity extraction model; if the entity extraction accuracy of the third entity extraction model exceeds a preset accuracy threshold, use the third entity extraction model as the target entity extraction model.
[0092] The second model training module 304 is also used to extract information from the second training data through the third entity extraction model to obtain second prediction data if the entity extraction accuracy of the third entity extraction model does not exceed the preset accuracy threshold; based on the data screening model, the second prediction data is screened according to the entity extraction template information and the entity extraction restriction information to obtain fourth training data; and incrementally train the third entity extraction model according to the fourth training data until the entity extraction accuracy of the third entity extraction model after the incremental training exceeds the preset accuracy threshold.
[0093] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0094] The method and apparatus of the present application can be used in many general or special computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer terminal devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices.
[0095] See also Figure 4 , Figure 4 A schematic block diagram of the structure of a network device provided in an embodiment of the present application.
[0096] like Figure 4 As shown, the network device 400 may include a processor 401 and a memory 402 , and the processor 401 and the memory 402 are connected via a bus 403 , such as an I2C (Inter-integrated Circuit) bus.
[0097] In an exemplary embodiment, the processor 401 can be used to provide computing and control capabilities to support the operation of the entire terminal device. The processor 401 can be a central processing unit (CPU), and the processor 401 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0098] Specifically, the memory 402 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a mobile hard disk.
[0099] Those skilled in the art will understand that Figure 4The structure shown in the figure is only a block diagram of a partial structure related to the embodiment of the present application, and does not constitute a limitation on the terminal device to which the embodiment of the present application is applied. The specific server may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.
[0100] Among them, the processor 401 is used to run the computer program stored in the memory, and implement any entity extraction model training method provided in the embodiments of the present application when executing the computer program.
[0101] In one embodiment, the processor 401 is used to run a computer program stored in a memory, and implement the following steps when executing the computer program: obtaining first training data and a first entity extraction model, and training the first entity extraction model according to the first training data to obtain a second entity extraction model; obtaining second training data, and extracting information from the second training data through the second entity extraction model to obtain first prediction data; determining entity extraction template information and entity extraction restriction information according to the second training data, and based on a data screening model, screening the first prediction data according to the entity extraction template information and the entity extraction restriction information to obtain third training data; and incrementally training the second entity extraction model according to the third training data to obtain a target entity extraction model.
[0102] In one embodiment, when implementing the method of determining entity extraction template information and entity extraction restriction information based on the second training data, the processor 401 is used to implement: determining the task type and task domain corresponding to the second training data; and determining the entity template information and entity restriction information based on the task type and task domain corresponding to the second training data.
[0103] In one embodiment, when the processor 401 implements the data screening model and filters the first prediction data according to the entity extraction template information and the entity extraction restriction information to obtain the third training data, it is used to implement: determining the data output template corresponding to the data screening model according to the entity template information; filtering the first prediction data according to the entity extraction restriction information, and integrating the filtered first prediction data based on the data output template to obtain the third training data.
[0104] In one embodiment, when the processor 401 implements the feature extraction of the target field information to obtain field feature information, it is used to implement: obtaining text data and segmenting the text data to obtain multiple segments of text information; determining the annotation format corresponding to each segment of the text information, and annotating the corresponding text information according to the annotation format corresponding to each segment of the text information to obtain the first training data.
[0105] In one embodiment, when the processor 401 implements the information extraction of the second training data through the second entity extraction model to obtain the first prediction data, it is used to implement: extracting information from the second training data through the second entity extraction model to obtain entity recognition information corresponding to the second training data; determining the confidence of each entity recognition information, and screening each entity recognition information according to the confidence of each entity recognition information to obtain the first prediction data.
[0106] In one embodiment, when the processor 401 implements the screening of each entity identification information according to the confidence of each entity identification information to obtain the first prediction data, it is used to implement: determining whether the confidence of each entity identification information exceeds a preset confidence; if the confidence of the entity identification information exceeds the preset confidence, then using the entity identification information as the first prediction data.
[0107] In one embodiment, when the processor 401 implements the incremental training of the second entity extraction model according to the third training data to obtain the target entity extraction model, it is used to implement: incremental training of the second entity extraction model according to the third training data to obtain a third entity extraction model; determining the entity extraction accuracy of the third entity extraction model; if the entity extraction accuracy of the third entity extraction model exceeds a preset accuracy threshold, using the third entity extraction model as the target entity extraction model. .
[0108] In one embodiment, after implementing the determination of the entity extraction accuracy of the third entity extraction model, the processor 401 is used to implement: if the entity extraction accuracy of the third entity extraction model does not exceed the preset accuracy threshold, extracting information from the second training data through the third entity extraction model to obtain second prediction data; based on the data screening model, screening the second prediction data according to the entity extraction template information and the entity extraction restriction information to obtain fourth training data; and incrementally training the third entity extraction model according to the fourth training data until the entity extraction accuracy of the third entity extraction model after the incremental training exceeds the preset accuracy threshold.
[0109] It should be noted that technicians in the relevant field can clearly understand that for the convenience and simplicity of description, the specific working process of the network device described above can refer to the corresponding process in the aforementioned terminal capability reporting method embodiment, and will not be repeated here.
[0110] An embodiment of the present application also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any entity extraction model training method provided in the description of the embodiment of the present application.
[0111] The storage medium may be an internal storage unit of the terminal device described in the foregoing embodiment, such as a hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc., equipped on the terminal device.
[0112] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0113] It should be understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system including the element.
[0114] The serial numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments. The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A training method for an entity extraction model, characterized in that: The method comprises: Acquire first training data and a first entity extraction model, and train the first entity extraction model according to the first training data to obtain a second entity extraction model; Acquire second training data, and extract information from the second training data using the second entity extraction model to obtain first prediction data; Determine entity extraction template information and entity extraction restriction information according to the second training data, and based on a data screening model, screen the first prediction data according to the entity extraction template information and the entity extraction restriction information to obtain third training data; Incrementally train the second entity extraction model according to the third training data to obtain a target entity extraction model.
2. The method according to claim 1, characterized in that The determining of entity extraction template information and entity extraction restriction information according to the second training data includes: Determining a task type and a task domain corresponding to the second training data; According to the task type and task domain corresponding to the second training data, the entity template information and the entity restriction information are determined.
3. The method according to claim 1, characterized in that The method of filtering the first prediction data based on the data screening model according to the entity extraction template information and the entity extraction restriction information to obtain the third training data includes: Determine a data output template corresponding to the data screening model according to the entity template information; The first prediction data is filtered according to the entity extraction restriction information, and the filtered first prediction data is integrated based on the data output template to obtain third training data.
4. The method according to claim 1, characterized in that: The obtaining of first training data comprises: Acquire text data, and segment the text data to obtain multiple segments of text information; Determine the annotation format corresponding to each segment of the text information, and annotate the corresponding text information according to the annotation format corresponding to each segment of the text information to obtain first training data.
5. The method according to claim 1, characterized in that The extracting information from the second training data by using the second entity extraction model to obtain first prediction data includes: Extracting information from the second training data using the second entity extraction model to obtain entity recognition information corresponding to the second training data; The confidence level of each entity identification information is determined, and each entity identification information is screened according to the confidence level of each entity identification information to obtain first prediction data.
6. The method according to claim 5, characterized in that The step of screening each entity identification information according to the confidence level of each entity identification information to obtain first prediction data includes: Determining whether the confidence level of each entity identification information exceeds a preset confidence level; If the confidence level of the entity recognition information exceeds a preset confidence level, the entity recognition information is used as the first prediction data.
7. The method according to claim 1, characterized in that The step of incrementally training the second entity extraction model according to the third training data to obtain a target entity extraction model includes: performing incremental training on the second entity extraction model according to the third training data to obtain a third entity extraction model; Determining the entity extraction accuracy of the third entity extraction model; If the entity extraction accuracy of the third entity extraction model exceeds a preset accuracy threshold, the third entity extraction model is used as the target entity extraction model.
8. The method according to claim 7, characterized in that After determining the entity extraction accuracy of the third entity extraction model, the method further includes: If the entity extraction accuracy of the third entity extraction model does not exceed the preset accuracy threshold, extracting information from the second training data through the third entity extraction model to obtain second prediction data; Based on the data screening model, the second prediction data is screened according to the entity extraction template information and the entity extraction restriction information to obtain fourth training data; Incremental training is performed on the third entity extraction model according to the fourth training data until the entity extraction accuracy of the third entity extraction model after the incremental training exceeds a preset accuracy threshold.
9. A network device, characterized in that: The network equipment includes: A processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing connection and communication between the processor and the memory, wherein when the computer program is executed by the processor, the steps of the training method of the entity extraction model as described in any one of claims 1 to 8 are realized.
10. A storage medium for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the training method of the entity extraction model as described in any one of claims 1 to 8.