Training File Generation and Evaluation Method, Device, Computer System, and Storage Medium
By working together with labeling servers, identification servers and hit servers, processing original files and calculating hit rate, the problem of not being able to know the hit rate of training samples in the prior art is solved, and the generation of high-quality training files and the rapid and accurate training of intelligent search models are achieved.
Patent Information
- Application Number
- CN202010344715.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-04-27
AI Technical Summary
The prior art cannot know the true hit rate of the training sample, which leads to the inability to ensure the labeling quality of the training sample, which in turn affects the fast and accurate training of the intelligent search model.
Through the collaborative work of the annotation server, the identification server and the hit server, the domain information and training entities of the original file are obtained, the original file is processed to generate the annotation file, the semantics of the annotation file are identified for sequence annotation, the intelligent search model is entered to obtain the training results, and the hit rate is calculated through the hit analysis algorithm to generate a hit analysis report.
Automatically generate training files, eliminate the impact of human errors, ensure the generation quality and speed of training files, and solve the problem of training sample labeling quality through hit rate evaluation, helping to quickly and accurately train the intelligent search model.
Smart Images

Figure CN111582497B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular, to a method, device, computer system, and storage medium for generating and evaluating training files. Background Art
[0002] A machine learning model is a general term for algorithms that can discover the implicit rules from a large amount of historical data to achieve prediction or classification. Specifically, it receives sample data and performs operations through its own functions to output prediction results or classification results. In the field of intelligent search, currently, sample files with annotations are usually used to train an intelligent search model based on a machine learning model to obtain a mature model that can accurately understand sample data and obtain accurate retrieval results based on this data.
[0003] Therefore, high-quality sample files are crucial for training an intelligent search model. However, since the current method for generating training files cannot know the true hit rate of training samples, the annotation quality of training samples cannot be guaranteed, resulting in the situation that an intelligent search model cannot be trained quickly and accurately. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, device, computer system, and storage medium for generating and evaluating training files, which are used to solve the problem in the prior art that the true hit rate of training samples cannot be known, resulting in the inability to guarantee the annotation quality of training samples.
[0005] To achieve the above purpose, the present invention provides a method for generating and evaluating training files, including:
[0006] A labeling server receives an original file and obtains the domain information and training entities of the original file, processes the original file according to the domain information and training entities to obtain a labeled file, and sends it to an identification server; wherein, the domain information is information data expressing the domain to which the original file belongs, and the training entity refers to the named entity in the original file.
[0007] The identification server identifies the semantics of the labeled file through a preset natural language understanding model, performs sequence labeling on it to obtain a training file, and sends the training file to a hit server.
[0008] The hit server has an intelligent search model and a hit analysis algorithm. The hit server inputs the training file into the intelligent search model corresponding to the domain information to obtain a training result, calculates the hit rate through the hit analysis algorithm, and summarizes the training file and the hit rate to generate a hit analysis report.
[0009] In the above solution, the steps of receiving the original file and obtaining the domain information and training entities of the original file include:
[0010] Obtain the original file, perform domain recognition on the original file to obtain domain information, and perform entity recognition on the original file to obtain independent entities;
[0011] Obtain the encoding of the independent entity through a preset relationship list and associate it with the independent entity;
[0012] Judge whether two adjacent independent entities have an associated relationship according to the preset relationship rules; if they have an associated relationship, merge the two independent entities to form an associated entity, and identify whether the next two adjacent independent entities have an associated relationship; if they do not have an associated relationship, identify whether the next two adjacent independent entities have an associated relationship;
[0013] Set the independent entity and the associated entity as training entities.
[0014] In the above solution, the steps of processing the original file according to the domain information and training entities to obtain an annotated file include:
[0015] Annotate the original file according to the training entities to obtain an annotated processing file;
[0016] Load the domain information into the annotated processing file to obtain an annotated file.
[0017] In the above solution, the steps of identifying the semantics of the annotated file and performing sequence annotation on it to obtain a training file include:
[0018] Perform semantic recognition on the annotated file to obtain a query intention;
[0019] Fill in the slot values of the annotated file according to the encoding in the annotated file to realize the sequence annotation of the training entities in the annotated file;
[0020] Summarize the query intention and the annotated file with sequence annotation to form a training file.
[0021] In the above solution, the steps of inputting the training file into the intelligent search model corresponding to the domain information to obtain a training result include:
[0022] Select the corresponding intelligent search model in the production environment according to the domain information of the training file, and input the training file into the intelligent search model;
[0023] The intelligent search model obtains a training result according to the query intention and the annotated file of the training file.
[0024] In the above solution, obtaining the hit rate by calculating the training result through the hit analysis algorithm includes:
[0025] Calculating the occurrence frequency of each training entity in the training file in the training result through the hit analysis algorithm to obtain the word frequency used to describe the importance degree of the training entity to the relevant file;
[0026] Calculating the quantity of each training entity in the training file in the training result through the hit analysis algorithm to obtain the inverse document frequency used to describe the scarcity degree of the training entity in the training result;
[0027] Multiplying the word frequency information and the inverse document frequency through the hit analysis algorithm to obtain the entity matching value used to describe the matching degree between each training entity and each relevant file;
[0028] Adding up the entity matching values of the relevant files to obtain the file matching value used to describe the matching degree between the training file and the relevant files;
[0029] Adding up the file matching values of each relevant file to obtain the hit rate used to describe the matching degree between the training file and the training result.
[0030] In the above solution, after summarizing the training file and the hit rate to generate a hit analysis report, it may further include:
[0031] Comparing the hit rate with a preset hit threshold;
[0032] If the hit rate exceeds the preset hit threshold, it is determined that the training file is qualified, and the hit analysis report is sent to the client;
[0033] If the hit rate does not exceed the preset hit threshold, it is determined that the training file is unqualified, and the hit analysis report is sent to the client.
[0034] To achieve the above object, the present invention further provides a training file generation and evaluation device, including:
[0035] A labeling server, configured to receive an original file and obtain the domain information and training entities of the original file, process the original file according to the domain information and training entities to obtain a labeled file, and send it to the recognition server; wherein, the domain information is information data expressing the domain to which the original file belongs, and the training entity refers to the named entity in the original file;
[0036] A recognition server, configured to recognize the semantics of the labeled file through a preset natural language understanding model, perform sequence labeling on it to obtain a training file, and send the training file to the hit server;
[0037] Hit the server, which has an intelligent search model and a hit analysis algorithm, to input the training file into the intelligent search model corresponding to the domain information to obtain a training result, calculate the hit rate of the training result through the hit analysis algorithm, and summarize the training file and the hit rate to generate a hit analysis report.
[0038] To achieve the above object, the present invention also provides a computer system, which includes a plurality of computer devices. Each computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processors of the plurality of computer devices execute the computer program, they jointly implement the steps of the above training file generation and evaluation method.
[0039] To achieve the above object, the present invention also provides a computer-readable storage medium, which includes a plurality of storage media. Each storage medium stores a computer program. When the computer programs stored on the plurality of storage media are executed by a processor, they jointly implement the steps of the above training file generation and evaluation method.
[0040] The training file generation and evaluation method, device, computer system, and storage medium provided by the present invention obtain an original file and obtain the domain information and training entities of the original file, process the original file according to the domain information and training entities to obtain an annotated file; and identify the semantics of the annotated file and perform sequence annotation on it to obtain a training file; so as to achieve the technical effect of automatically obtaining the training file, eliminating the influence of human errors, and ensuring the generation quality and generation speed of the training file.
[0041] Input the training file into the intelligent search model corresponding to the domain information to obtain a training result, calculate the hit rate of the training result through the hit analysis algorithm, and summarize the training file and the hit rate to generate a hit analysis report. Therefore, by sending the hit rate of the training result to the user side, the problem that the true hit rate of the training sample cannot be known currently, resulting in the annotation quality of the training sample not being guaranteed, is solved. Brief Description of the Drawings
[0042] Figure 1 It is a flowchart of the first embodiment of the training file generation and evaluation method of the present invention;
[0043] Figure 2 It is a flowchart of obtaining a training data set in the first embodiment of the training file generation and evaluation method of the present invention;
[0044] Figure 3 It is a flowchart of obtaining an annotated file in the first embodiment of the training file generation and evaluation method of the present invention;
[0045] Figure 4Flowchart for obtaining a training file in Embodiment 1 of the method for generating and evaluating a training file of the present invention;
[0046] Figure 5 Flowchart for obtaining a training result in Embodiment 1 of the method for generating and evaluating a training file of the present invention;
[0047] Figure 6 Flowchart for obtaining a hit rate in Embodiment 1 of the method for generating and evaluating a training file of the present invention;
[0048] Figure 7 Flowchart after generating a hit analysis report in Embodiment 1 of the method for generating and evaluating a training file of the present invention;
[0049] Figure 8 Schematic diagram of program modules in Embodiment 2 of the apparatus for generating and evaluating a training file of the present invention;
[0050] Figure 9 Schematic diagram of the hardware structure of a computer device in Embodiment 3 of the computer system of the present invention.
[0051] Reference numerals:
[0052] 1. Apparatus for generating and evaluating a training file 2. Computer device 11. Annotation server
[0053] 12. Recognition server 13. Hit server 21. Memory 22. Processor Detailed implementation manners
[0054] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0055] The training file generation and evaluation method, device, computer system, and storage medium provided by the present invention are applicable to the field of machine learning, and aim to provide a training file generation and evaluation method based on an annotation server, an identification server, and a hit server. The present invention receives an original file through the annotation server and obtains the domain information and training entities of the original file, processes the original file according to the domain information and training entities to obtain an annotation file; the identification server identifies the semantics of the annotation file through a preset natural language understanding model and performs sequence annotation on it to obtain a training file; the hit server has an intelligent search model and a hit analysis algorithm, the hit server inputs the training file into the intelligent search model corresponding to the domain information to obtain a training result, calculates the training result through the hit analysis algorithm to obtain a hit rate, and summarizes the training file and the hit rate to generate a hit analysis report.
[0056] Embodiment 1
[0057] Please refer to Figure 1 , a training file generation and evaluation method in this embodiment includes:
[0058] S1: The annotation server receives the original file and obtains the domain information and training entities of the original file, processes the original file according to the domain information and training entities to obtain an annotation file, and sends it to the identification server; wherein, the domain information is information data expressing the domain to which the original file belongs, and the training entity refers to the named entity in the original file;
[0059] S2: The identification server identifies the semantics of the annotation file through a preset natural language understanding model and performs sequence annotation on it to obtain a training file, and sends the training file to the hit server;
[0060] S3: The hit server has an intelligent search model and a hit analysis algorithm, the hit server inputs the training file into the intelligent search model corresponding to the domain information to obtain a training result, calculates the training result through the hit analysis algorithm to obtain a hit rate, and summarizes the training file and the hit rate to generate a hit analysis report.
[0061] In this application, the original document can be an article or a short sentence stored in a database, or a query term or query statement output by a client. In this embodiment, the domain information can be fund audit, or intelligent supervision, or macro decision-making; the annotation document refers to the text information obtained by annotating the original document according to training entities. The semantic recognition of the annotation document is performed through a natural language understanding model to obtain the query intention of the annotation document; the sequence annotation of the annotation document is performed through the natural language understanding model. In this embodiment, the sequence annotation of the annotation document is performed by the method of slot value filling; the query intention is loaded into the annotation document with sequence annotation to obtain a training document.
[0062] The hit rate algorithm uses the TF-IDF (Term Frequency Inverse Document Frequency) algorithm, which is a commonly used weighting algorithm for information retrieval and text mining. It is used to evaluate the importance of a word for a document set or a single document in a corpus. Among them, the importance of a word increases in direct proportion to the number of times it appears in the document, but at the same time decreases in inverse proportion to the frequency of its appearance in the corpus.
[0063] Therefore, the training document generation and evaluation method provided by the present invention obtains a training document by obtaining an original document and obtaining the domain information and training entities of the original document, processing the original document according to the domain information and training entities to obtain an annotation document; and identifying the semantics of the annotation document and performing sequence annotation on it to obtain a training document; to achieve the technical effect of automatically obtaining a training document, eliminating the influence of human errors, and ensuring the generation quality and generation speed of the training document.
[0064] The training document is input into an intelligent search model corresponding to the domain information to obtain a training result, the hit rate is calculated by the hit analysis algorithm for the training result, and the training document and the hit rate are summarized to generate a hit analysis report; therefore, by sending the hit rate of the training result to the client, it provides an index reference for the annotation management entity recognition model and / or the natural language understanding model, so as to help the annotation management personnel obtain high-quality sample documents, and is used to achieve the technical effect of quickly and accurately training the intelligent search model, solving the problem that the true hit rate of the training sample cannot be known currently, resulting in the annotation quality of the training sample not being guaranteed.
[0065] In a preferred embodiment, please refer to Figure 2 , the receiving the original document and obtaining the domain information and training entities of the original document includes:
[0066] S101: Obtain the original document, perform domain recognition on the original document to obtain domain information, and perform entity recognition on the original document to obtain independent entities;
[0067] In this step, the original file is obtained by extracting it from a storage server that pre-stores the original file, or by receiving the original file output by the client.
[0068] S102: Obtain the code of the independent entity through a preset relationship list and associate it with the independent entity.
[0069] In this step, the relationship list exists in the relational database. The relationship list includes codes and code entities. Among them, the code corresponds to at least one code entity. Obtain the code entity corresponding to the independent entity in the relationship list, set the code corresponding to the code entity as the target code, and associate it with the independent entity by loading the target code into the original file. For example: Assume the codes include date and location, that is: DATE and LOCATION; the
[0070] code entities corresponding to the date code include: yesterday, today, tomorrow; the code entities corresponding to the location code include: Beijing, Shanghai, Guangzhou, Shenzhen; if the original file is: What's the weather like in Shenzhen today, then the independent entity associated with the code is as follows:
[0071] Today Shenzhen
[0072] DATE LOCATION
[0073] Among them, DATE refers to the date code, and LOCATION refers to the location code.
[0074] S103: Judge whether two adjacent independent entities have an association relationship according to the preset relationship rules; if they have an association relationship, merge the two independent entities to form an associated entity, and identify whether the next two adjacent independent entities have an association relationship; if they do not have an association relationship, identify whether the next two adjacent independent entities have an association relationship.
[0075] In this step, the relationship rules are used to stipulate the association relationship between codes. Extract the codes of two adjacent independent entities in the original file, and judge whether there is an association relationship between the two codes according to the relationship rules; if there is an association relationship, copy the independent entities corresponding to the two codes and merge them to form an associated entity.
[0076] For example: The original document is the 2019 Medical Insurance Policy of Jiangsu Province. In step S101, two independent entities, namely Jiangsu and Medical Insurance Policy, are identified. In step S102, the code "LOCATION" for "Jiangsu" and the code "POLICY" for "Medical Insurance Policy" are obtained. If there is an associated relationship between the code "LOCATION" and the code "POLICY" in the relationship rule, then the associated entity of the "Medical Insurance Policy of Jiangsu Province" will be obtained.
[0077] S104: Set the independent entity and the associated entity as training entities.
[0078] In this step, the independent entity and the associated entity are set as training entities and duplicate items are removed to ensure the brevity and accuracy of the training entities.
[0079] In a preferred embodiment, the obtaining of the domain information by performing domain recognition on the original document and the obtaining of the independent entity by performing entity recognition on the original document include:
[0080] S101-1: Identify the original document through a preset domain list to obtain the words corresponding to the domain keywords in the domain list.
[0081] In this step, the domain list includes a domain title and domain keywords, and there is at least one domain keyword under each domain title. In this embodiment, the domain titles at least include Fund Audit, Intelligent Supervision, and Macro Decision-making.
[0082] S101-2: Obtain the number of occurrences of the words in the original document, set the domain keyword corresponding to the word with the largest number of occurrences as the target keyword, and set the domain title of the target keyword in the domain list as the domain information.
[0083] In this step, in the original document, the number of occurrences of the words corresponding to the domain keywords is obtained in sequence, the word with the largest number of occurrences is obtained, and the domain keyword corresponding to this word is set as the target keyword.
[0084] S101-3: Perform entity recognition on the words in the original document through an entity recognition model to obtain the independent entities in the original document.
[0085] In this embodiment, the entity recognition model is a Conditional Random Field (CRF) model. The conditional random field model is a discriminative probability model, which is a type of random field and is commonly used for labeling or analyzing sequence data, such as natural language text or biological sequences. Since it is common knowledge in the art for those skilled in the art to obtain named entities through the conditional random field model, and the technical problem to be solved in this step is how to obtain the domain to which the original document belongs and its independent entities, the specific process of the conditional random field model will not be elaborated in this application.
[0086] In a preferred embodiment, after entity recognition of the original document to obtain independent entities, the following steps are further included:
[0087] S101-4: Identify the independent entity through a synonym database prestored with synonyms, obtain synonyms having the same meaning as the independent entity, and set the synonyms as the independent entity.
[0088] In this step, the synonym database has a synonym set, and the synonyms in the synonym set have the same meaning. The independent entity is compared with the synonym set in sequence to obtain words consistent with the independent entity, and all the synonyms in the synonym set where the words are located are set as the independent entity.
[0089] In a preferred embodiment, please refer to Figure 3 , the process of processing the original document according to the domain information and training entities to obtain an annotated document includes:
[0090] S111: Annotate the original document according to the training entity to obtain an annotated processing document;
[0091] In this step, the words in the original document are annotated according to the independent entities and associated entities in the training entity to obtain an annotated processing document.
[0092] S112: Load the domain information into the annotated processing document to obtain an annotated document.
[0093] In this step, the domain information is used as part of the title of the annotated processing document or part of the file name of the annotated processing document to achieve the effect of loading the domain information into the annotated processing document. At this time, the annotated processing document will be converted into an annotated document.
[0094] In a preferred embodiment, please refer to Figure 4 , the process of identifying the semantics of the annotated document and performing sequence annotation on it to obtain a training document includes:
[0095] S201: Perform semantic recognition on the annotated document to obtain a query intention.
[0096] In this step, the semantic recognition is essentially a text classification task, which identifies the semantics of the annotated document through a natural language understanding model to obtain the query intention of the annotated document; among them, the original document has at least one query intention. For example: The original document is: "What's the weather like in Shenzhen today?", at this time, what the user expresses is to query the weather, and here we can consider querying the weather as an intention.
[0097] It should be noted that the natural language understanding (NLU) is a computer algorithm for semantic recognition of text. Since those skilled in the art can easily identify the text semantics through the natural language understanding model, and the problem to be solved by this application is how to know whether the training file meets the operator's expectations, the working principle of the natural language understanding model will not be elaborated here.
[0098] S202: Fill the slot values of the annotation file according to the encoding in the annotation file to achieve sequence annotation of the training entities in the annotation file.
[0099] In this step, through the natural language understanding model and according to the encoding in the annotation file, the slot values of the annotation file are filled, so that the model can accurately perform sequence annotation on the independent entities or associated entities with encoding in the annotation file.
[0100] In this embodiment, the slot value filling is essentially a task of performing sequence annotation on the entities in the annotation file in the form of BIO.
[0101] Based on the above example, for instance, the original file is: "What's the weather like in Shenzhen today?", at this time what the user expresses is to query the weather. Here we can consider querying the weather as an intention. Then specifically which place's weather and which day's weather are being queried. Here the user also conveys this information, (location = Shenzhen, date = today). And here the independent entities or associated entities corresponding to the location encoding and date encoding are information slots. Still taking "What's the weather like in Shenzhen today?" as an example, when performing intention recognition, it is classified into the intention of "asking about the weather" by the text classification method, and when performing slot value filling, it can be annotated as follows by the sequence annotation method:
[0102] What's the weather like in Shenzhen today
[0103] B_DATE B_LOCATION O OOOOO.
[0104] It should be noted that slot value filling is a task of performing sequence annotation on the entities in the text based on natural language understanding technology, which belongs to the prior art. Therefore, those skilled in the art can easily perform sequence annotation on the text through natural language understanding technology. And the problem to be solved by this application is how to perform sequence annotation on the valuable entities in the text in a targeted manner. Therefore, the specific process of slot value filling will not be elaborated in this application.
[0105] S203: Summarize the query intention and the annotation file with sequence annotation to form a training file.
[0106] In a preferred embodiment, please refer to Figure 5, the step of inputting the training file into the intelligent search model corresponding to the domain information to obtain a training result includes:
[0107] S301: Select a corresponding intelligent search model in the production environment according to the domain information of the training file, and input the training file into the intelligent search model;
[0108] In this step, the intelligent search model in the production environment has professional labels, and the professional labels are used to describe the fields that the intelligent search model is good at predicting or classifying; obtain the professional label that matches the domain information in the production environment, select the intelligent search model corresponding to the professional label as the target model, and input the training file into the target model.
[0109] It should be noted that the production environment refers to the service system that officially provides external services. The intelligent search model refers to a search engine built based on a machine learning model and set in the server of the service system; the machine learning model is a general term for algorithms used to mine the implicit rules from a large amount of historical data and used for prediction or classification; the intelligent search model receives sample data and performs operations through its own function to output a prediction result or a classification result.
[0110] S302: The intelligent search model obtains a training result according to the query intention and annotation file of the training file.
[0111] Among them, the training result refers to the prediction result or classification result obtained by the intelligent search model through its own function to calculate the training file.
[0112] In a preferred embodiment, please refer to Figure 6 , the step of calculating the hit rate by the hit analysis algorithm to obtain the hit rate includes:
[0113] S311: Calculate the occurrence frequency of each training entity in the training result in the training file by the hit analysis algorithm to obtain the word frequency used to describe the importance of the training entity to the relevant file.
[0114] In this step, the word frequency refers to the frequency of a certain word (Term) appearing in a document. In this embodiment, frequency is used instead of the number of times, and the purpose is to prevent the situation that some words appear too many times due to the too long document content.
[0115] In this embodiment, the hit analysis algorithm has a word frequency objective function, and the word frequency of each training entity in the training result is calculated through the word frequency objective function.
[0116] Among them, the word frequency objective function is:
[0117]
[0118] In the above formula, tfi,j refers to the word frequency of the i-th training entity in the training file in the j-th relevant file; ni,j refers to the i-th training entity in the training file and the number of occurrences of the i-th training entity in the j-th relevant file in the training result. The denominator ∑knk,j refers to the sum of the number of occurrences of all training entities (where the training entities are k in number) in the training file in the relevant file.
[0119] Through the above method, the normalization processing of each training entity in the training file is realized, and the importance degree of each training entity to each relevant file is correctly evaluated, that is: the importance degree of a training entity in a relevant file increases as the number of occurrences of the training entity increases.
[0120] For example: the total number of words in a certain relevant file in the training result is 100, and the word "Shanghai" appears 3 times. Then the word frequency of the word "Shanghai" in this file is 3 / 100 = 0.03.
[0121] S312: Calculate the number of each training entity in the training result in the training file through the hit analysis algorithm to obtain the inverse document frequency for describing the scarcity degree of the training entity in the training result.
[0122] In this step, the inverse document frequency (IDF) refers to the number of documents containing a certain word in a document set. It represents the general importance degree of a training entity in the training result.
[0123] In this embodiment, the hit analysis algorithm includes an inverse objective function, and the number of each training entity in the training result is calculated through the inverse objective function.
[0124] Among them, the inverse objective function is as follows:
[0125]
[0126] In the above formula, idfi refers to the inverse document frequency of the i-th training entity in the training file in the training result, |D| represents the total number of files in the document set, that is, the total number of relevant files in the training result of this application; |{j:ti∈dj}| refers to the number of relevant files containing the word ti (that is, the number of files with ni≠0); therefore, the inverse document frequency represents the importance of a training entity in a training result. The rarer it is, the higher the weight, so it decreases as the number of words increases. Based on the above example, if the training entity "Shanghai" appears in 1,000 relevant files, and the total number of relevant files in the training result is 10,000,000, then the inverse document frequency of the training entity "Shanghai" is log(10,000,000 / 1,000) = 4.
[0127] Optionally, since a training entity may not be in the training result, once such a training entity is encountered, the inverse objective function will be incorrect or the function will fail because its denominator is zero, which may further cause errors or even crashes in the computer program; therefore, by adding a natural number to the denominator of the inverse objective function, it is ensured that the denominator will never be zero under any circumstances, thus avoiding the situation where the inverse objective function is incorrect or the function fails. For example, adding the natural number 1 to the denominator makes the denominator expressed as follows:
[0128] 1 + |{d∈D:f∈d}|
[0129] S313: Multiply the word frequency information and the inverse document frequency through the hit analysis algorithm to obtain an entity matching value for describing the matching degree between each training entity and each relevant file.
[0130] In this step, the hit analysis algorithm has a hit objective function, and the entity matching value between each training entity and each relevant file is obtained through the hit objective function; where
[0131] The hit objective function is shown as follows:
[0132] tfidf i,j = tf i,j × idf i
[0133] Among them, tfidfi,j refers to the entity matching value between the i-th training entity and the j-th relevant file. tfi,j refers to the word frequency of the i-th training entity in the training file in the j-th relevant file. idfi refers to the inverse document frequency of the i-th training entity in the training file in the training result. In summary, the hit objective function can produce a high entity matching value tf-idf for a high training entity frequency within a certain relevant file and a low document frequency of the training entity in the entire training result. Therefore, the hit objective function tends to filter out common words and retain important words.
[0134] For example, based on the above example, the obtained entity matching value is: tfidfi,j = 0.03 × 4 = 0.12.
[0135] S314: Add up the entity matching values of the relevant file to obtain a file matching value for describing the matching degree between the training file and the relevant file.
[0136] In this step, by adding up the entity matching degrees between all training entities in the training file and a certain relevant file, the matching degree between the training file and the relevant file, as well as the file matching value describing this matching degree, can be obtained.
[0137] Based on the above example, if the training file includes training entities: Beijing, Shanghai, Guangzhou, Shenzhen; if the entity matching value of Beijing with the j-th relevant file is: 0.03; if the entity matching value of Beijing with the j-th relevant file is: 0.12; if the entity matching value of Beijing with the j-th relevant file is: 0.01; if the entity matching value of Beijing with the j-th relevant file is: 0.10; then the file matching value of this training file with the j-th relevant file is: 0.25.
[0138] S315: Add up the file matching values of each relevant file to obtain a hit rate for describing the matching degree between the training file and the training result.
[0139] In this step, the hit rate can be obtained by adding up the file matching values of all relevant files in the training result, or the training result can be sorted in descending order according to the file matching value, and the file matching values in the top (such as the top ten) can be added up to obtain the hit rate.
[0140] In a preferred embodiment, please refer to Figure 7 , after summarizing the training file and the hit rate to generate a hit analysis report, it may further include:
[0141] S321: Compare the hit rate with a preset hit threshold;
[0142] S322: If the hit rate exceeds a preset hit threshold, determine that the training file is qualified and send the hit analysis report to the client;
[0143] S323: If the hit rate does not exceed the preset hit threshold, determine that the training file is unqualified and send the hit analysis report to the client.
[0144] In this step, if the hit rate exceeds the preset hit threshold, it means that the domain information and annotation file obtained through the original file
[0145] and the accuracy of semantic recognition and sequence annotation of the annotation file meet the requirements, and the corresponding training file is qualified.
[0146] If the hit rate does not exceed the preset hit threshold, it means that the domain information and annotation file obtained through the original file, and the accuracy of semantic recognition and sequence annotation of the annotation file do not meet the requirements, and the corresponding training file is unqualified. Therefore, the annotator can use the hit analysis report as a reference for the adjustment relationship list, and / or domain list, and / or relationship rules, and / or synonym database, and / or entity recognition model, and / or natural language understanding model to obtain a training file with a hit rate exceeding the hit threshold.
[0147] Embodiment 2
[0148] Please refer to Figure 8 , a training file generation and evaluation device 1 of this embodiment includes:
[0149] An annotation server 11, configured to receive an original file and obtain the domain information and training entities of the original file, process the original file according to the domain information and training entities to obtain an annotation file, and send it to the recognition server 12; wherein, the domain information is information data expressing the domain to which the original file belongs, and the training entity refers to the named entity in the original file;
[0150] A recognition server 12, configured to recognize the semantics of the annotation file through a preset natural language understanding model, perform sequence annotation on it to obtain a training file, and send the training file to the hit server 13;
[0151] A hit server 13, having an intelligent search model and a hit analysis algorithm, configured to input the training file into the intelligent search model corresponding to the domain information to obtain a training result, calculate the training result through the hit analysis algorithm to obtain a hit rate, and summarize the training file and the hit rate to generate a hit analysis report.
[0152] The present technical solution can be applied to the field of model hosting in artificial intelligence. By obtaining the domain information and training entity of the original file, processing the original file according to the domain information and training entity to obtain an annotated file, identifying the semantics of the annotated file, and performing sequence annotation on it to obtain a training file, inputting the training file into an intelligent search model corresponding to the domain information to obtain a training result, calculating the training result through a hit analysis algorithm to obtain a hit rate, and summarizing the training file and the hit rate to generate a hit analysis report, it realizes improving the generation quality and generation speed of the training file, and provides an index reference for the annotation management entity recognition model and / or natural language understanding model, so as to help the annotation management personnel obtain high-quality sample files, and further help with the machine learning tasks in the model construction process.
[0153] Embodiment III:
[0154] To achieve the above object, the present invention further provides a computer system. The computer system includes a plurality of computer devices 2. The components of the training file generation and evaluation device 1 in Embodiment II can be dispersed in different computer devices. The computer device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a rack-mounted server, a blade server, a tower server or a cabinet server (including an independent server, or a server cluster composed of multiple servers) that executes a program, etc. The computer device in this embodiment at least includes, but is not limited to: a memory 21 and a processor 22 that can communicate with each other through a system bus, as Figure 9 shown. It should be noted that Figure 9 only a computer device with components - is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0155] In this embodiment, the memory 21 (i.e., the readable storage medium) includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 21 may be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory 21 may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device. Of course, the memory 21 may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory 21 is generally used to store the operating system and various application software installed on the computer device, such as the program code of the training file generation and evaluation device in Embodiment 1. In addition, the memory 21 may also be used to temporarily store various data that have been output or will be output.
[0156] In some embodiments, the processor 22 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 22 is generally used to control the overall operation of the computer device. In this embodiment, the processor 22 is used to run the program code stored in the memory 21 or process data, such as running the training file generation and evaluation device to implement the training file generation and evaluation method in Embodiment 1.
[0157] Embodiment 4:
[0158] To achieve the above object, the present invention also provides a computer-readable storage system, which includes a plurality of storage media, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, server, App application mall, etc., on which a computer program is stored, and when the program is executed by the processor 22, the corresponding functions are implemented. The computer-readable storage medium of this embodiment is used to store the training file generation and evaluation device, and when it is executed by the processor 22, the training file generation and evaluation method in Embodiment 1 is implemented.
[0159] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0161] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for generating and evaluating training files, characterized in that, it includes: The annotation server receives the original file and obtains the domain information and training entities of the original file, processes the original file according to the domain information and training entities to obtain an annotation file, and sends it to the recognition server; wherein, the domain information is information data expressing the domain to which the original file belongs, and the training entity refers to the named entity in the original file; The recognition server recognizes the semantics of the annotation file through a preset natural language understanding model, performs sequence annotation on it to obtain a training file, and sends the training file to the hit server; The hit server has an intelligent search model and a hit analysis algorithm. The hit server inputs the training file into the intelligent search model corresponding to the domain information to obtain a training result, calculates the training result through the hit analysis algorithm to obtain a hit rate, and summarizes the training file and the hit rate to generate a hit analysis report; The receiving the original file and obtaining the domain information and training entities of the original file includes: Obtaining the original file, performing domain recognition on the original file to obtain domain information, and performing entity recognition on the original file to obtain independent entities; Obtaining the encoding of the independent entity through a preset relationship list, and associating it with the independent entity; Judging whether two adjacent independent entities have an association relationship according to a preset relationship rule; if they have an association relationship, merging the two independent entities to form an associated entity, and identifying whether the next two adjacent independent entities have an association relationship; if they do not have an association relationship, identifying whether the next two adjacent independent entities have an association relationship; Setting the independent entity and the associated entity as training entities; The performing domain recognition on the original file to obtain domain information and performing entity recognition on the original file to obtain independent entities includes: Identifying the original file through a preset domain list to obtain words corresponding to the domain keywords in the domain list; wherein, the domain list includes a domain title and domain keywords, and there is at least one domain keyword under each domain title; Obtaining the number of occurrences of the word in the original file, setting the domain keyword corresponding to the word with the largest number as the target keyword, and setting the domain title of the target keyword in the domain list as the domain information; Performing entity recognition on the words in the original file through an entity recognition model to obtain independent entities in the original file; wherein, the entity recognition model is a conditional random field model.
2. The method for generating and evaluating training files according to claim 1, characterized in that, The processing the original file according to the domain information and training entities to obtain an annotation file includes: Annotating the original file according to the training entity to obtain an annotated processing file; Loading the domain information into the annotated processing file to obtain an annotation file.
3. The method for generating and evaluating training files according to claim 1, characterized in that, The recognizing the semantics of the annotation file and performing sequence annotation on it to obtain a training file includes: Performing semantic recognition on the annotation file to obtain a query intention; Perform slot value filling on the annotation file according to the encoding in the annotation file to achieve sequence annotation of the training entities in the annotation file; Summarize the query intent and the annotation file with sequence annotation to form a training file.
4. The training file generation and evaluation method according to claim 1, characterized in that, the step of inputting the training file into the intelligent search model corresponding to the domain information to obtain a training result includes: Select a corresponding intelligent search model in the production environment according to the domain information of the training file, and input the training file into the intelligent search model; The intelligent search model obtains a training result according to the query intent and annotation file of the training file.
5. The training file generation and evaluation method according to claim 1, characterized in that, the step of calculating the hit rate by the hit analysis algorithm for the training result includes: Calculate the occurrence frequency of each training entity in the training file in the training result through the hit analysis algorithm to obtain the word frequency used to describe the importance of the training entity to the relevant file; Calculate the quantity of each training entity in the training file in the training result through the hit analysis algorithm to obtain the inverse document frequency used to describe the scarcity degree of the training entity in the training result; Multiply the word frequency information and the inverse document frequency through the hit analysis algorithm to obtain an entity matching value used to describe the matching degree between each training entity and each relevant file; Add up the entity matching values of the relevant files to obtain a file matching value used to describe the matching degree between the training file and the relevant files; Add up the file matching values of each relevant file to obtain the hit rate used to describe the matching degree between the training file and the training result.
6. The training file generation and evaluation method according to claim 1, characterized in that, after summarizing the training file and the hit rate to generate a hit analysis report, it may further include: Compare the hit rate with a preset hit threshold; If the hit rate exceeds the preset hit threshold, determine that the training file is qualified and send the hit analysis report to the client; If the hit rate does not exceed the preset hit threshold, determine that the training file is unqualified and send the hit analysis report to the client.
7. A training file generation and evaluation device, characterized in that, comprising: A labeling server, configured to receive an original file and obtain domain information and training entities of the original file, process the original file according to the domain information and training entities to obtain a labeled file, and send the labeled file to an identification server; wherein, the domain information is information data expressing the domain to which the original file belongs, and the training entity refers to a named entity in the original file; the receiving the original file and obtaining the domain information and training entities of the original file includes: obtaining the original file, performing domain recognition on the original file to obtain domain information, and performing entity recognition on the original file to obtain independent entities; obtaining encodings of the independent entities through a preset relationship list and associating them with the independent entities; judging whether two adjacent independent entities have an associated relationship according to a preset relationship rule; if they have an associated relationship, merging the two independent entities to form an associated entity, and identifying whether the next two adjacent independent entities have an associated relationship; if they do not have an associated relationship, identifying whether the next two adjacent independent entities have an associated relationship; setting the independent entities and associated entities as training entities; the performing domain recognition on the original file to obtain domain information and performing entity recognition on the original file to obtain independent entities includes: recognizing the original file through a preset domain list to obtain words corresponding to domain keywords in the domain list; wherein, the domain list includes domain titles and domain keywords, and each domain title has at least one domain keyword; obtaining the number of occurrences of the words in the original file, setting the domain keyword corresponding to the word with the largest number as the target keyword, and setting the domain title of the target keyword in the domain list as the domain information; performing entity recognition on the words in the original file through an entity recognition model to obtain independent entities in the original file; wherein, the entity recognition model is a conditional random field model; An identification server, configured to recognize the semantics of the labeled file through a preset natural language understanding model, perform sequence labeling on the labeled file to obtain a training file, and send the training file to a hit server; A hit server, having an intelligent search model and a hit analysis algorithm, configured to input the training file into the intelligent search model corresponding to the domain information to obtain a training result, calculate the training result through the hit analysis algorithm to obtain a hit rate, and summarize the training file and the hit rate to generate a hit analysis report.
8. A computer system, which includes multiple computer devices, and each computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processors of the multiple computer devices execute the computer program, they jointly implement the steps of the training file generation and evaluation method according to any one of claims 1 to 6.
9. A computer-readable storage medium, which includes multiple storage media, and each storage medium stores a computer program, wherein, When the computer programs stored in the multiple storage media are executed by a processor, they jointly implement the steps of the training file generation and evaluation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Named entity relation extraction and construction method based on deep learning
CN104199972A
Training method and system of named entity recognition model and electronic device
CN109190110A
Search method and device, computer equipment and storage medium
CN110765275A