Sample data processing method, device, computer program product and storage medium
By screening sample texts based on the support and confusion of syntax structure and entity types in the sequence labeling model training, the problems of low sample selection efficiency and uneven sample in the prior art are solved, and the model performance is improved.
Patent Information
- Application Number
- CN202111417183.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-25
AI Technical Summary
The prior art is low in selecting samples trained by sequence annotation model, ignoring samples with lower probability, and failing to consider the similarity of syntactic structures between samples, resulting in uneven samples and affecting model performance.
By obtaining the pending text set, the syntactic structure of the sample text and its proportion of the number in the text set, and the sample text is input into the pre-trained named entity recognition model, the type tag of the entity and its proportion of the number in the text set, the support and confusion of the sample text are calculated, and the target sample text is selected from the text set based on these indicators.
Through the threshold filtering of support and confusion, sample texts with similar syntactic structures are selected to avoid sample imbalance. At the same time, the amount of information in the sample text is limited by confusion, and sample texts with higher value are selected to improve the performance of the sequence labeling model.
Smart Images

Figure CN114219012B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence and natural language processing, and in particular to a method, device, computer program product and storage medium for processing sample data. Background Art
[0002] In the field of natural language processing, sequence labeling is the main task at the sentence level, which is used to predict the labels that need to be annotated in a given text sequence.
[0003] In order to improve the performance of the sequence labeling model, high-quality samples are needed to train the sequence labeling model. Samples with high confusion and large information content can play a greater role in the optimization process of the sequence labeling model. Therefore, how to select more valuable samples is an important factor in improving the performance of the sequence labeling model.
[0004] In related technologies, in order to select the samples that are "most likely to be confused" or have the largest amount of "information", the following methods are usually used: the least confidence selection method (Least Confident), which can select the sample with the highest predicted probability but the lowest "credibility"; the minimum margin sampling method (margin sampling), which can select the sample with the smallest difference between the two probability values predicted by the model. The least confidence selection method ignores the samples with lower probabilities, and the minimum margin sampling method also only considers the two samples with the highest predicted probabilities, resulting in low efficiency in sample selection. Summary of the invention
[0005] The embodiments of the present disclosure provide a method, apparatus, computer program product, and storage medium for processing sample data to select sample texts with higher value.
[0006] One aspect of an embodiment of the present disclosure provides a method for processing sample data, including: obtaining a text set to be processed, the text set including more than one sample text; determining the syntactic structure of the sample text, and determining the quantitative proportion of the syntactic structure in the text set; inputting the sample text into a pre-trained named entity recognition model, and outputting the boundary labels of the words in the sample text and the confidence of the boundary labels through the named entity recognition model; determining the entities included in the sample text based on the boundary labels, and determining the F value and type label of the entity; determining the quantitative proportion of the type label of the entity in the text set; determining the support of the sample text based on the quantitative proportion of the type label of the entity in the text set, the F value of the entity, and the quantitative proportion of the syntactic structure in the text set; determining the confusion of the sample text based on the confidence of the boundary label; and selecting a target sample text from the text set based on the support, confusion, preset support threshold and confusion threshold of the sample text.
[0007] In some embodiments, the confusion degree of the sample text is determined based on the confidence of the boundary label, including: determining the information entropy of the word based on the confidence of the boundary label; determining the mean of the information entropy of the words included in the entity as the confusion degree of the entity; determining the mean of the confusion degrees of the entities included in the sample text as the confusion degree of the sample text.
[0008] In some embodiments, the method also includes determining the confidence of the type label of the entity; determining the information entropy of the word based on the confidence of the boundary label, including: adjusting the confidence of the boundary label based on the confidence of the type label of the entity to which the word belongs to obtain the adjusted confidence of the boundary label; determining the information entropy of the word based on the adjusted confidence of the boundary label.
[0009] In some embodiments, the support of the sample text is positively correlated with the quantitative proportion of the syntactic structure in the text set, the support of the sample text is positively correlated with a first numerical value, and the support of the sample text is negatively correlated with a second numerical value, wherein the first numerical value is the average quantitative proportion of the type labels of each entity included in the sample in the text set, and the second numerical value is the average F value of each entity included in the sample text.
[0010] In some embodiments, the syntactic structure is determined via the following steps: segmenting the sample text to obtain a segmentation sequence; determining the part of speech of the words in the segmentation sequence; and performing syntactic analysis on the segmentation sequence based on the part of speech to obtain a syntactic structure.
[0011] In some embodiments, the type label of the entity is obtained through the following steps: using Elastic Search to perform entity type recall on the entity to obtain the type label of the entity.
[0012] In some embodiments, the method further includes: constructing a sample set based on the target sample text; and training a pre-constructed initial sequence labeling model based on the sample set to obtain a trained sequence labeling model.
[0013] According to another aspect of the embodiments of the present disclosure, there is provided an apparatus for processing sample data, comprising: a data acquisition unit configured to acquire a text set to be processed, the text set comprising more than one sample text; a syntax determination unit configured to determine the syntactic structure of the sample text and determine the quantitative proportion of the syntactic structure in the text set; a text recognition unit configured to input the sample text into a pre-trained named entity recognition model, and output the boundary labels of the characters included in the sample text and the confidence of the boundary labels through the named entity recognition model; an entity determination unit configured to determine the entities included in the sample text based on the boundary labels, and determine the F value and type label of the entity; a type proportion unit configured to determine the quantitative proportion of the type label of the entity in the text set; a support unit configured to determine the support of the sample text based on the quantitative proportion of the type label of the entity in the text set, the F value of the entity and the quantitative proportion of the syntactic structure in the text set; a confusion unit configured to determine the confusion of the sample text based on the confidence of the boundary label; and a sample selection unit configured to select a target sample text from the text set based on a preset support threshold and confusion threshold.
[0014] Another aspect of the embodiments of the present disclosure provides a computer program product, including a computer program / instruction, which implements the method for processing sample data in any of the above embodiments when executed by a processor.
[0015] Another aspect of the embodiments of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for processing sample data in any of the above embodiments is implemented.
[0016] The method for processing sample data provided by the embodiment of the present disclosure first determines the syntactic structure of the sample text and determines the proportion of the syntactic structure in the text set; then the sample text is input into a pre-trained named entity recognition model to determine the boundary labels of the characters in the sample text and the confidence of the boundary labels; then, based on the boundary labels, the entities included in the sample text are determined, and the F value and type label of the entity are determined; then, the proportion of the type label of the entity in the text set is determined; then, based on the proportion of the type label of the entity in the text set, the F value of the entity and the proportion of the syntactic structure in the text set, the support of the sample text is determined; and, based on the confidence of the boundary label, the confusion of the sample text is determined; finally, based on the preset support threshold and confusion threshold, the target sample text is selected from the text set. By limiting the number of sample texts with similar syntactic structures through support, the sample imbalance caused by too many samples of the same syntactic structure is avoided, and at the same time, the amount of information in the sample text is limited by confusion, so that a sample text with higher value can be selected.
[0017] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0019] The present disclosure may be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:
[0020] Figure 1 A flowchart of an embodiment of a method for processing sample data disclosed herein;
[0021] Figure 2 A flowchart for determining the degree of confusion in one embodiment of sample data processing of the present disclosure;
[0022] Figure 3 A flowchart for determining a syntactic structure in one embodiment of a method for sample data processing of the present disclosure;
[0023] Figure 4 A flowchart of another embodiment of the method for processing sample data of the present disclosure;
[0024] Figure 5 A schematic diagram of the structure of an embodiment of a device for processing sample data disclosed in the present invention;
[0025] Figure 6 The figure is a schematic diagram of the structure of an application embodiment of the electronic device disclosed in the present invention. DETAILED DESCRIPTION
[0026] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless otherwise specifically stated.
[0027] Those skilled in the art can understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate the necessary logical order between them.
[0028] It should also be understood that in the embodiments of the present disclosure, “plurality” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.
[0029] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.
[0030] In addition, the term "and / or" in the present disclosure is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present disclosure generally indicates that the associated objects before and after are in an "or" relationship.
[0031] It should also be understood that the description of the various embodiments in the present disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced to each other, and for the sake of brevity, they will not be described one by one.
[0032] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0033] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0034] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.
[0035] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0036] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate with many other general or special computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, small computer systems, large computer systems, and distributed cloud computing technology environments including any of the above systems, etc.
[0037] Electronic devices such as terminal devices, computer systems, servers, etc. can be described in the general context of computer system executable instructions (such as program modules) executed by computer systems. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0038] SUMMARY OF THE DISCLOSURE
[0039] In the process of implementing the present disclosure, the inventors found that the minimum confidence selection method ignored the samples with lower probability, and the minimum spacing sample selection method also only considered the two samples with the largest prediction probability, resulting in low efficiency of sample selection. In addition, the above two methods do not consider the similarity of syntactic structure between samples. When the number of similar samples is large, it will lead to sample imbalance, thereby affecting the performance of the sequence labeling model.
[0040] Exemplary Methods
[0041] Next reference Figure 1 , Figure 1 A flow chart showing an embodiment of a method for processing sample data of the present disclosure is shown as follows: Figure 1 As shown, the process includes the following steps.
[0042] Step 110: Obtain a text set to be processed.
[0043] The text set includes more than one sample text.
[0044] As an example, the execution subject may be, for example, a terminal device or a server, which may obtain the text set to be processed through a network.
[0045] Step 120: Determine the syntactic structure of the sample text, and determine the percentage of the syntactic structure in the text set.
[0046] The syntactic structure represents the combination of words. For example, the combination may include verb-object structure, subject-predicate structure, attributive-predicate structure, and other types.
[0047] In this embodiment, the proportion of the number of syntactic structures in the text set represents the ratio of the number of sample texts with the same syntactic structure to the total number of sample texts.
[0048] As an example, the execution entity may perform syntactic analysis on the sample text by using the HanLP tool to determine the syntactic structure of the sample text.
[0049] Step 130: Input the sample text into a pre-trained named entity recognition model, and the named entity recognition model outputs the boundary labels of the characters in the sample text and the confidence of the boundary labels.
[0050] As an example, the bilstm+crf model can be used to perform entity recognition on the sample text to determine the boundary label of each word in the sample text and the confidence of each boundary label. For example, the boundary labels can include three types: B, I, and O, where B represents the starting point of the entity boundary, I represents an entity, and O represents a non-entity. The sum of the confidences of the three labels corresponding to the same word is 1.
[0051] Step 140: determine the entities included in the sample text based on the boundary labels, and determine the F value and type label of the entity.
[0052] Usually, the F value (i.e., F-score) is used to characterize the performance of an entity and can be calculated by the precision and recall of the entity.
[0053] In this embodiment, the type tag of the entity is used to characterize the type of the entity. For example, the entity tag of "apple" can be "fruit", "mobile phone" and other types.
[0054] In a specific example, the execution subject can take the word with the boundary label "B" as the starting point of the interval, and then take the word after the word and before the first word with the boundary label "O" as the end point of the interval, and extract the words in the interval in order to obtain the entity composed of words. Afterwards, the execution subject can determine the F value corresponding to the entity according to the precision and recall rate corresponding to the entity, and use the entity recognition model (such as the baseline model) to recognize the entity and determine the type label of the entity.
[0055] In some optional implementations of this embodiment, the type label of the entity is obtained through the following steps: using Elastic Search to recall the entity type to obtain the type label of the entity.
[0056] In this implementation, the execution subject can input the index value corresponding to the entity (for example, the Chinese pinyin of the entity) into Elastic Search to obtain the type label corresponding to the index value, which can improve the efficiency of determining the entity type.
[0057] As an example, the execution entity can retrieve "pingguo" through Elastic Search, and the recalled entities may include "apple, mobile phone", "apple, fruit" arranged by confidence, where "mobile phone" and "fruit" are type labels of the entity "apple".
[0058] Step 150: Determine the percentage of entity type labels in the text set.
[0059] In this embodiment, the proportion of the type labels of entities in the text set represents the ratio of the number of entities with the same type labels to the total number of entities in the text set.
[0060] Step 160: Determine the support of the sample text based on the proportion of the entity type label in the text set, the F value of the entity, and the proportion of the syntactic structure in the text set.
[0061] In this embodiment, the support of the sample texts may represent the similarity of the syntactic structures of the sample texts.
[0062] In some optional implementations of this embodiment, the support of the sample text is positively correlated with the proportion of the syntactic structure in the text set, the support of the sample text is positively correlated with a first numerical value, and the support of the sample text is negatively correlated with a second numerical value, wherein the first numerical value is the mean of the proportion of the type labels of each entity included in the sample in the text set, and the second numerical value is the mean of the F values of each entity included in the sample text.
[0063] As an example, the support of the sample text can be determined by the following formulas (1), (2), and (3).
[0064]
[0065]
[0066]
[0067] In the formula, s_sample represents the support of the sample text, s_sentence represents the proportion of the syntactic structure in the text set, s_entity represents the first value, and s_entity j represents the proportion of the type label of the jth entity in the text set, a represents the first entity in the sample text, b represents the last entity in the sample text, m represents the number of entities in the sample text, and F j represents the F value of the jth entity, and f represents the second numerical value.
[0068] Step 170: Determine the confusion degree of the sample text based on the confidence of the boundary label.
[0069] As an example, the execution subject may first determine the information entropy of each word based on the confidence of the boundary label, and then determine the average of the information entropy of each word as the confusion degree of the sample text. For example, the information entropy of a word may be determined by the following formula (4):
[0070] Hi =-(P B *log2 P B +P I *log2 P I +P O *log2 P O ) (4)
[0071] In the formula, H i represents the information entropy of the i-th word, P B represents the confidence of the boundary label "B", P I represents the confidence of the boundary label “I”, P O Indicates the confidence of the boundary label "O".
[0072] Step 180: Obtain target sample text from the text set based on the support, confusion, preset support threshold and confusion threshold of the sample text.
[0073] As an example, the support threshold can be set to 0.8 and the confusion threshold can be set to 0.7 based on experience. The execution body can traverse the sample texts in the text set and then determine the sample texts with support greater than 0.8 and confusion greater than 0.7 as target sample texts.
[0074] The method for processing sample data provided by the embodiment of the present disclosure first determines the syntactic structure of the sample text and determines the proportion of the syntactic structure in the text set; then the sample text is input into a pre-trained named entity recognition model to determine the boundary labels of the characters in the sample text and the confidence of the boundary labels; then, based on the boundary labels, the entities included in the sample text are determined, and the F value and type label of the entity are determined; then, the proportion of the type label of the entity in the text set is determined; then, based on the proportion of the type label of the entity in the text set, the F value of the entity and the proportion of the syntactic structure in the text set, the support of the sample text is determined; and, based on the confidence of the boundary label, the confusion of the sample text is determined; finally, based on the preset support threshold and confusion threshold, the target sample text is selected from the text set. By limiting the number of sample texts with similar syntactic structures through support, the sample imbalance caused by too many samples of the same syntactic structure is avoided, and at the same time, the amount of information in the sample text is limited by confusion, so that a sample text with higher value can be selected.
[0075] Next reference Figure 2 , Figure 2 A flowchart of determining the degree of confusion in one embodiment of sample data processing of the present disclosure is shown. Figure 1 In some optional implementations of the embodiment shown, step 170 may also be implemented by Figure 2 The process shown includes the following steps.
[0076] Step 210: Determine the information entropy of the word based on the confidence of the boundary label.
[0077] Step 220: Determine the mean value of the information entropy of the characters included in the entity as the confusion degree of the entity.
[0078] In this embodiment, the execution entity may determine the mean value of the information entropy of the characters included in the entity as the confusion degree of the entity.
[0079] Step 230: Determine the average of the confusion degrees of the entities included in the sample text as the confusion degree of the sample text.
[0080] In this embodiment, the execution entity may determine the average of the confusion degrees of the entities included in the sample text as the confusion degree of the sample text.
[0081] In a specific example, the execution subject can determine the confusion degree of the entity through formula (5) and determine the confusion degree of the sample text through formula (6).
[0082]
[0083]
[0084] In the formula, H_entity j represents the confusion degree of the jth entity, H_sample represents the confusion degree of the sample text, l represents the end position of the entity, k represents the start position of the entity, n represents the number of characters in the entity, m represents the number of entities in the sample text, a represents the first entity in the sample text, and b represents the last entity in the sample text.
[0085] Figure 2 The process of determining the confusion degree of the sample text shown can determine the confusion degree of the entity based on the information entropy of the character, and then determine the confusion degree of the sample text based on the confusion degree of the entity, thereby realizing the calculation of the confusion degree for the multi-label classification task, which can more accurately reflect the amount of information of the sample.
[0086] In some optional implementations of this embodiment, the method also includes determining the confidence of the type label of the entity; the above step 210 may also include: adjusting the confidence of the boundary label according to the confidence of the type label of the entity to which the character belongs, to obtain the adjusted confidence of the boundary label; based on the adjusted confidence of the boundary label, determining the information entropy of the character.
[0087] As an example, the execution subject may use Elastic Search to determine the confidence of the type label of the entity, then multiply the confidence of the entity type label by the confidence of the boundary label of each word included in the entity, and determine the obtained product as the adjusted boundary confidence of the word, and then determine the information entropy of the word based on the adjusted boundary confidence.
[0088] For example, the execution subject can adjust the confidence of the boundary label through the following formulas (7), (8), and (9):
[0089] P B =P_ES j *p B (7)
[0090] P I =P_ES j *p I (8)
[0091] P o =P_ES j *p O (9)
[0092] In the formula, p B represents the confidence of the boundary label B, P B represents the confidence of the adjusted boundary label B, p I represents the confidence of the boundary label I, P I represents the confidence of the adjusted boundary label I, p O represents the confidence of the boundary label O, P O represents the confidence of the adjusted boundary label O, P_ES j Represents the confidence of the type label of the jth entity.
[0093] In this implementation, the confidence of the entity type label can be introduced into the calculation process of the character's information entropy, so that the character's information entropy can include the entity's type information. The confusion degree of the sample text thus obtained can reflect the amount of information about the entity type contained in the sample text, which can improve the value of the target sample text for the entity recognition model.
[0094] Next reference Figure 3 , Figure 3 A flowchart of determining a syntactic structure in an embodiment of a method for processing sample data of the present disclosure is shown. Figure 3 As shown, the process includes the following steps.
[0095] Step 310: Segment the sample text into words to obtain a word segmentation sequence.
[0096] Step 320: Determine the part of speech of the words in the word segmentation sequence.
[0097] For example, parts of speech may include adjectives, verbs, nouns, and the like.
[0098] Step 330: Based on the part of speech, perform syntactic analysis on the word segmentation sequence to obtain a syntactic structure.
[0099] In a specific example, the sample text is "I bought a house in Xi'erqi". After the execution subject performs word segmentation on the sample text, the analysis sequence obtained is: "I", "bought", "Xi'erqi", "house", and then the part of speech of each word is determined respectively, for example, the part of speech of "I" is a pronoun, and the part of speech of "house" is a noun. After that, the execution subject can use HanLP to perform syntactic analysis on the word segmentation sequence, and the obtained syntactic structure may include: subject-predicate relationship, verb-object relationship, attributive-predicate relationship, right-attached relationship and core relationship.
[0100] In this embodiment, the execution subject can segment the sample text and determine the part of speech of the segmented words, and then perform syntactic analysis on the segmented word sequence according to the part of speech to determine the syntactic structure of the sample text, which can improve the accuracy and efficiency of the syntactic structure.
[0101] Next reference Figure 4 , Figure 4 FIG. 1 is a flow chart showing another embodiment of the method for processing sample data of the present disclosure. Figure 4 As shown, in Figure 1 Based on the process shown, the method may also include Figure 4 The process shown includes the following steps.
[0102] Step 410: construct a sample set based on the target sample text.
[0103] Step 420: Based on the sample set, train the pre-built initial sequence labeling model to obtain a trained sequence labeling model.
[0104] In this embodiment, the target sample text selected based on confusion and support can have a large amount of information, and the target sample text in the sample set is evenly distributed, which can improve the training quality of the sequence labeling model and help improve the performance of the sequence labeling model.
[0105] Exemplary Devices
[0106] Next reference Figure 5 , Figure 5 A schematic diagram showing a structure of an embodiment of a device for processing sample data of the present disclosure is shown as follows: Figure 5As shown, the device includes: a data acquisition unit 510, configured to acquire a text set to be processed, the text set including more than one sample text; a syntax determination unit 520, configured to determine the syntactic structure of the sample text, and determine the proportion of the syntactic structure in the text set; a text recognition unit 530, configured to input the sample text into a pre-trained named entity recognition model, and the named entity recognition model outputs the boundary labels of the words included in the sample text and the confidence of the boundary labels; an entity determination unit 540, configured to determine the entities included in the sample text based on the boundary labels, and determine the F value and type label; a type proportion unit 550 is configured to determine the quantitative proportion of the entity's type label in the text set; a support unit 560 is configured to determine the support of the sample text based on the quantitative proportion of the entity's type label in the text set, the entity's F value and the quantitative proportion of the syntactic structure in the text set; a confusion unit 570 is configured to determine the confusion of the sample text based on the confidence of the boundary label; a sample selection unit 580 is configured to obtain a target sample text from the text set based on the support, confusion, preset support threshold and confusion threshold of the sample text.
[0107] In this embodiment, the confusion unit 570 further includes: an information entropy module, configured to determine the information entropy of a word based on the confidence of a boundary label; an entity confusion module, configured to determine the mean of the information entropy of the words included in the entity as the confusion degree of the entity; and a sample confusion module, configured to determine the mean of the confusion degrees of the entities included in the sample text as the confusion degree of the sample text.
[0108] In this embodiment, the device also includes a type confidence unit, which is configured to determine the confidence of the type label of the entity; the information entropy module further includes: an adjustment submodule, which is configured to adjust the confidence of the boundary label based on the confidence of the type label of the entity to which the word belongs, and obtain the adjusted confidence of the boundary label; an information entropy determination submodule, which is configured to determine the information entropy of the word based on the adjusted confidence of the boundary label.
[0109] In this embodiment, the support of the sample text is positively correlated with the quantitative proportion of the syntactic structure in the text set, the support of the sample text is positively correlated with the first numerical value, and the support of the sample text is negatively correlated with the second numerical value, wherein the first numerical value is the average quantitative proportion of the type labels of each entity included in the sample in the text set, and the second numerical value is the average F value of each entity included in the sample text.
[0110] In this embodiment, the syntax determination unit 520 further includes: a word segmentation module, configured to perform word segmentation on the sample text to obtain a word segmentation sequence; a part-of-speech recognition module, configured to determine the part-of-speech of the words in the word segmentation sequence; and a structural analysis module, configured to perform syntactic analysis on the word segmentation sequence based on the part-of-speech to obtain a syntactic structure.
[0111] In this embodiment, the entity determination unit 540 is further configured to: use Elastic Search to recall entity types for entities to obtain type labels for the entities.
[0112] In this embodiment, the device also includes: a sample construction unit, configured to construct a sample set based on the target sample text; a training unit, configured to train a pre-constructed initial sequence labeling model based on the sample set to obtain a trained sequence labeling model.
[0113] In addition, an embodiment of the present disclosure further provides an electronic device, including:
[0114] Memory for storing computer programs;
[0115] The processor is used to execute the computer program stored in the memory, and when the computer program is executed, the method for processing sample data described in any of the above embodiments of the present disclosure is implemented.
[0116] Figure 6 This is a schematic diagram of the structure of an application embodiment of the electronic device disclosed in the present invention. Figure 6 The electronic device according to the embodiment of the present disclosure is described. The electronic device may be any one or both of the first device and the second device, or a stand-alone device independent of them, and the stand-alone device may communicate with the first device and the second device to receive the collected input signals from them.
[0117] like Figure 6 As shown, the electronic device includes one or more processors and memory.
[0118] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0119] The memory may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, a random access memory (RAM) and / or a cache memory (cache), etc. The non-volatile memory may include, for example, a read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute the program instructions to implement the sample data processing method of each embodiment of the present disclosure described above and / or other desired functions.
[0120] In one example, the electronic device may further include: an input device and an output device, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0121] In addition, the input device may also include, for example, a keyboard, a mouse, and the like.
[0122] The output device can output various information to the outside, including the determined distance information, direction information, etc. The output device can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.
[0123] Of course, to simplify, Figure 6 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application situations, the electronic device may further include any other appropriate components.
[0124] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps in the method of sample data processing according to various embodiments of the present disclosure described in the above part of this specification.
[0125] The computer program product may be written in any combination of one or more programming languages to write program code for performing the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0126] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enables the processor to execute the steps of the method for processing sample data according to various embodiments of the present disclosure described in the above part of this specification.
[0127] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0128] A person of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk, etc., various media that can store program codes.
[0129] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0130] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0131] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," and the like are open words, referring to "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or," and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0132] The method and apparatus of the present disclosure may be implemented in many ways. For example, the method and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, which includes machine-readable instructions for implementing the method according to the present disclosure. Therefore, the present disclosure also covers a recording medium storing a program for executing the method according to the present disclosure.
[0133] It should also be noted that in the apparatus, device and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0134] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0135] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A sample data processing method, characterized in that: include: Obtaining a text set to be processed, wherein the text set includes more than one sample text; Determine the syntactic structure of the sample text, and determine the proportion of the syntactic structure in the text set; Inputting the sample text into a pre-trained named entity recognition model, and outputting boundary labels of words included in the sample text and confidences of the boundary labels through the named entity recognition model; Determine the entity included in the sample text based on the boundary label, and determine the F value and type label of the entity; Determine the proportion of the type label of the entity in the text set; Determining the support of the sample text based on the proportion of the type label of the entity in the text set, the F value of the entity, and the proportion of the syntactic structure in the text set; Determining the confusion degree of the sample text based on the confidence of the boundary label; Based on the support, confusion, preset support threshold and confusion threshold of the sample text, a target sample text is obtained from the text set.
2. The method according to claim 1, characterized in that Determining the confusion degree of the sample text based on the confidence of the boundary label includes: Determining the information entropy of the word based on the confidence of the boundary label; Determine the mean value of information entropy of the words included in the entity as the confusion degree of the entity; The average of the confusion degrees of the entities included in the sample text is determined as the confusion degree of the sample text.
3. The method according to claim 2, characterized in that The method also includes determining a confidence level of the type label of the entity; Determining the information entropy of the word based on the confidence of the boundary label includes: Adjusting the confidence of the boundary label based on the confidence of the type label of the entity to which the character belongs, to obtain an adjusted confidence of the boundary label; Based on the confidence of the adjusted boundary label, the information entropy of the word is determined.
4. The method according to any one of claims 1 to 3, characterized in that The support of the sample text is positively correlated with the proportion of the syntactic structure in the text set, the support of the sample text is positively correlated with a first numerical value, and the support of the sample text is negatively correlated with a second numerical value, wherein the first numerical value is the average proportion of the type labels of each of the entities included in the sample in the text set, and the second numerical value is the average F value of each of the entities included in the sample text.
5. The method according to any one of claims 1 to 3, characterized in that The syntactic structure is determined via the following steps: Segmenting the sample text to obtain a segmentation sequence; Determining the part of speech of the words in the word segmentation sequence; Based on the parts of speech, the word segmentation sequence is subjected to syntactic analysis to obtain the syntactic structure.
6. The method according to any one of claims 1 to 3, characterized in that The type label of the entity is obtained through the following steps: Elastic Search is used to recall the entity type of the entity to obtain the type label of the entity.
7. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Based on the target sample text, construct a sample set; Based on the sample set, a pre-built initial sequence labeling model is trained to obtain a trained sequence labeling model.
8. A device for processing sample data, characterized in that: include: A data acquisition unit is configured to acquire a text set to be processed, wherein the text set includes more than one sample text; A syntax determination unit, configured to determine the syntax structure of the sample text and determine the percentage of the syntax structure in the text set; A text recognition unit is configured to input the sample text into a pre-trained named entity recognition model, and output boundary labels of words included in the sample text and confidences of the boundary labels through the named entity recognition model; An entity determination unit, configured to determine an entity included in the sample text based on the boundary label, and determine an F value and a type label of the entity; A type proportion unit, configured to determine a quantitative proportion of the type label of the entity in the text set; A support unit is configured to determine the support of the sample text based on the proportion of the type label of the entity in the text set, the F value of the entity, and the proportion of the syntactic structure in the text set; A confusion degree unit, configured to determine the confusion degree of the sample text based on the confidence of the boundary label; The sample selection unit is configured to obtain a target sample text from the text set based on the support, confusion, preset support threshold and confusion threshold of the sample text.
9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Text processing method and related device
CN112287111A
Word Segmentation method and System for Language Text
US20190018836A1