Corpus redundancy removal method, device, computer equipment and storage medium
By identifying and removing redundant sentences in the corpus, the problems of reduced diversity and low training efficiency caused by corpus redundancy in the prior art are solved, and more efficient model training is achieved.
Patent Information
- Application Number
- CN202210404831.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-04-18
AI Technical Summary
When removing corpus redundancy, the prior art can easily weaken the diversity of model training corpus, increase the redundancy of repeated corpus, and affect the model training effect.
By obtaining the pending corpus, determining the category of participle entity in the sentence, determining the semantic characteristics of the sentence based on the participle category, performing clustering processing to identify redundant clustering, and removing sentences in redundant clustering.
Effectively remove corpus redundancy, maintain corpus diversity, improve model training effect, and accelerate the model training process.
Smart Images

Figure CN114840664B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, computer equipment and storage medium for removing corpus redundancy. Background Art
[0002] Machine learning requires a large amount of training data, especially entity recognition in natural language processing, which requires as diverse training corpus as possible. As the training corpus increases, corpus redundancy is inevitable. Corpus that cannot provide effective information has no positive significance for entity recognition models. This type of corpus is considered redundant. Deleting such corpus can further optimize model performance and accelerate model training. Removing redundancy before annotation can also reduce annotation pressure.
[0003] However, a common method to remove redundant corpus is to use a language model to score the sentences in the corpus. The higher the score, the closer the sentence is to the field of interest, while the lower the score, the farther the sentence is from the field of interest. Sentences with too low scores are removed as redundant. However, the problem here is that it tends to find a large amount of similar corpus, that is, if a sentence scores very high, its similar sentences will also score very high, which in turn reduces the diversity of the model training corpus and increases the redundancy of repeated corpus. Summary of the invention
[0004] In order to solve the above technical problems, the present application provides a method, apparatus, computer device and storage medium for removing corpus redundancy.
[0005] In a first aspect, the present application provides a method for removing corpus redundancy, comprising:
[0006] Acquire a corpus to be processed, wherein the corpus to be processed includes a plurality of sentences, and each of the sentences includes a plurality of participles;
[0007] Determining the entity category of each of the participles in each of the sentences;
[0008] Determining the semantic features of each of the sentences based on the entity categories of all the participles in each of the sentences;
[0009] Performing clustering processing on the semantic features of all the sentences in the corpus to be processed to determine redundant clusters, wherein the semantic features in the redundant clusters are used to indicate the sentences with repeated semantics;
[0010] The sentences corresponding to each of the semantic features in the redundant clusters are removed.
[0011] In a second aspect, the present application provides a corpus redundancy removal device, comprising:
[0012] An acquisition module, used for acquiring a corpus to be processed, wherein the corpus to be processed includes a plurality of sentences, each of which includes a plurality of participles;
[0013] A recognition module, used to determine the entity category of each of the word segments in each of the sentences;
[0014] A determination module, configured to determine the semantic features of each of the sentences based on the entity categories of all the segmented words in each of the sentences;
[0015] A clustering module, used for clustering the semantic features of all the sentences in the corpus to be processed to determine redundant clusters, wherein the semantic features in the redundant clusters are used to indicate the sentences with repeated semantics;
[0016] A removal module is used to remove the sentences corresponding to each of the semantic features in the redundant cluster.
[0017] In a third aspect, the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:
[0018] Acquire a corpus to be processed, wherein the corpus to be processed includes a plurality of sentences, and each of the sentences includes a plurality of participles;
[0019] Determining the entity category of each of the participles in each of the sentences;
[0020] Determining the semantic features of each of the sentences based on the entity categories of all the participles in each of the sentences;
[0021] Performing clustering processing on the semantic features of all the sentences in the corpus to be processed to determine redundant clusters, wherein the semantic features in the redundant clusters are used to indicate the sentences with repeated semantics;
[0022] The sentences corresponding to each of the semantic features in the redundant clusters are removed.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0024] Acquire a corpus to be processed, wherein the corpus to be processed includes a plurality of sentences, and each of the sentences includes a plurality of participles;
[0025] Determining the entity category of each of the participles in each of the sentences;
[0026] Determining the semantic features of each of the sentences based on the entity categories of all the participles in each of the sentences;
[0027] Performing clustering processing on the semantic features of all the sentences in the corpus to be processed to determine redundant clusters, wherein the semantic features in the redundant clusters are used to indicate the sentences with repeated semantics;
[0028] The sentences corresponding to each of the semantic features in the redundant clusters are removed.
[0029] The above-mentioned corpus redundancy removal method is applied to the field of deep learning technology to realize natural language processing. Based on the above-mentioned corpus redundancy removal method, the entity category of each word in each sentence in the acquired corpus to be processed is determined, and the semantic features of the sentence are determined by using the entity categories of all the words in the sentence. The semantic features of all the sentences in the corpus to be processed are clustered to obtain redundant clusters. The semantic features in the redundant clusters are used to indicate the existence of the sentences with repeated semantics. The sentences corresponding to each semantic feature in the redundant clusters are removed to complete the removal of redundant corpus. Removing similar corpus not only reduces the training corpus, but also ensures the diversity of the corpus, which can improve the subsequent model training effect and accelerate model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0032] Figure 1 1 is a flow chart of a method for removing corpus redundancy in one embodiment;
[0033] Figure 2 is a structural block diagram of a corpus redundancy removal device in one embodiment;
[0034] Figure 3 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0035] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0036] In one embodiment, Figure 1 FIG. 1 is a flow chart of a method for removing redundant corpus in an embodiment, referring to FIG. Figure 1 , a method for removing redundant corpus is provided. This embodiment mainly uses the method applied to a server as an example to illustrate that the method for removing redundant corpus specifically includes the following steps:
[0037] Step S110, obtaining the corpus to be processed.
[0038] The corpus to be processed includes a plurality of sentences, and each of the sentences includes a plurality of participles.
[0039] Specifically, the corpus to be processed is used as training data for natural language machine learning. The word segments in the corpus to be processed can be text characters corresponding to any language, such as Chinese characters, English characters, etc. In this embodiment, Chinese characters are selected to form the corpus to be processed.
[0040] Step S120, determining the entity category of each of the word segments in each of the sentences.
[0041] Specifically, the entity category is used to indicate semantic categories such as segmentation, which can be names of people, places, institutions, time expressions, book titles, song titles, etc. Specifically, the entity recognition model can be used to perform category recognition processing on each segmentation in the corpus to be processed to output the entity category corresponding to each segmentation. The entity recognition model can be a hidden Markov model (HMM), a maximum entropy model (MEM), a support vector machine (SVM), a conditional random field (CRF), a NN / CNN-CRF deep learning model, a RNN-CRF deep learning model, etc.
[0042] Step S130, determining the semantic features of each of the sentences based on the entity categories of all the word segments in each of the sentences.
[0043] Specifically, the semantic features of each sentence in the corpus to be processed are used to indicate the distribution of corresponding participles of different entity categories in the sentence, that is, there may be participles corresponding to multiple different entity categories in the sentence, there may also be participles corresponding to the same entity category, and there may also be participles whose entity categories are confused, that is, the same participle may correspond to multiple different entity categories. For example, sub-participles corresponding to institution names appear in participles corresponding to place names. Assume that the entity category corresponding to the participle "Maotai Town" is a place name, but the entity category corresponding to "Maotai" is an institution name, that is, the entity category of the participle is confused.
[0044] Step S140: clustering the semantic features of all the sentences in the corpus to be processed to determine redundant clusters.
[0045] The semantic features in the redundant cluster are used to indicate the sentences with repeated semantics.
[0046] Specifically, the semantic features corresponding to all sentences in the processed corpus are clustered, that is, the semantic features are classified to distinguish valid semantic features from invalid semantic features. Valid semantic features are used to indicate sentences without repeated entity contexts and confusion in the segmentation entity categories. Conversely, invalid semantic features are used to indicate sentences with repeated entity contexts and / or no confusion in the segmentation entity categories. That is, invalid semantic features are clustered to form redundant clusters. Confusion in the segmentation entity category indicates that there are multiple ways of understanding the segmentation, which can be used to increase the diversity of training data. No confusion in the segmentation entity category indicates that there is only one way to understand the segmentation. Sentences without entity category confusion are judged as redundant corpus with a single corpus and meaningless repetitions.
[0047] Step S150: removing the sentences corresponding to the respective semantic features in the redundant clusters.
[0048] Specifically, sentences corresponding to each semantic feature in the redundant cluster are removed to complete the removal of redundant corpus. Removing similar corpus not only reduces the training corpus, but also ensures the diversity of the corpus, which can improve the subsequent model training effect and accelerate model training. Performing the above-mentioned corpus redundancy removal operation before the annotation work begins can also reduce the annotation pressure of the annotators.
[0049] In one embodiment, the determining of the semantic features of each sentence based on the entity categories of all the segmentations in each sentence includes: counting entity confusion counts, the number of segmentations under each entity category, and the total number of segmentations under all entity categories according to the entity categories of all the segmentations in the sentence; and determining the semantic features of the sentence according to the entity confusion counts, the number of segmentations under each entity category, and the total number of segmentations under all entity categories.
[0050] The entity confusion count is used to indicate the frequency of the presence of a second participle under the second category in a first participle of the first category in the sentence, the first category and the second category are different entity categories, the first participle and the second participle are both participles in the sentence, and the length of the first participle is greater than the length of the second participle.
[0051] Specifically, the entity categories of each participle in each sentence are summarized and counted to determine the number of participles corresponding to each entity category. In this embodiment, the entity categories are names of people, places, organization names, and time expressions. Specifically, the number of participles corresponding to names is recorded as P, the number of participles corresponding to place names is recorded as L, the number of participles corresponding to organization names is recorded as O, and the number of participles corresponding to time expressions is recorded as T. Then the total number of participles corresponding to all entity categories is counted, and the total number of participles is recorded as S.
[0052] The entity confusion count is used to indicate the frequency of the confused participle in the sentence. The confused participle is used to indicate the participle corresponding to multiple entity categories. The confused participle is the first participle corresponding to the first category mentioned above. There are also sub-participles corresponding to other entity categories in the confused participle. The sub-participle here is the second participle under the second category mentioned above. For example, the confused participle is "Liu Chongqing". The first category corresponding to the confused participle is a person's name, but the second category corresponding to the sub-participle "Chongqing" in the confused participle is a place name. Therefore, when counting the entity confusion count, the participle "Liu Chongqing" is counted once. There may also be multiple sub-participles corresponding to different second categories in the confused participle.
[0053] When the first category is a person's name, the second category can be at least one of a place name, an organization name, or a time expression. The segmentation confusion count obtained when the first category is a person's name is recorded as P2; when the first category is a place name and the second category is at least one of a person's name, an organization name, or a time expression, the segmentation confusion count obtained is recorded as L2; when the first category is an organization name and the second category is at least one of a person's name, a place name, or a time expression, the segmentation confusion count obtained is recorded as O2. In this way, based on the entity category of each segmentation and the nested relationship between each segmentation, the confusing segmentations in each sentence are counted, the segmentation confusion counts corresponding to each entity category are determined, and then the segmentation confusion counts corresponding to each entity list are accumulated to obtain the entity confusion count.
[0054] Since confused segmentation corresponds to multiple entity categories, that is, there are multiple ways to understand confused segmentation, it can be used to enhance the diversity of subsequent model learning. The semantic features of the sentence can be determined by comprehensively considering the entity confusion count, the number of segmentations under each entity category, and the total number of segmentations under all entity categories. The diversity of sentence understanding can be determined through the semantic features.
[0055] In one embodiment, counting entity confusion counts according to entity categories of all the participles in the sentence includes: when there are multiple second participles of the second category in the first participle, accumulating the entity confusion count of the first participle once.
[0056] Specifically, in order to avoid repeated counting when counting entity confusion counts and affecting the semantic features of the sentence, no matter how many second participles corresponding to the second category exist in the first participle, the first participle is only counted once when counting entity confusion counts. For example, the first participle is "Three Points Yiwu Company", and the first category corresponding to the first participle is the name of the organization, wherein the second category corresponding to the second participle "Three Points" is a time expression, and the second category corresponding to the second participle "Yiwu" is a place name, but the confusion count in the first participle is only counted once, that is, O2 is accumulated by one. In this way, the participle confusion counts corresponding to each entity category are determined, and then it is determined that the entity confusion counts obtained after the participle confusion counts corresponding to each entity category are aggregated do not have repeated counts.
[0057] In one embodiment, the semantic features include a prominence ratio and a confusion ratio, and determining the semantic features of the sentence according to the entity confusion count, the number of participles under each entity category, and the total number of participles under all entity categories includes: obtaining the prominence ratio according to the ratio of the number of participles with the largest value to the total number of participles among the numbers of participles under each entity category; and obtaining the confusion ratio according to the ratio of the entity confusion count to the total number of participles.
[0058] Specifically, the prominence ratio is denoted as R1, R1 = max(P, L, O, T) / S, that is, among the number of segmentations corresponding to each entity category, the ratio of the number of segmentations with the largest value to the total number of segmentations is taken as the prominence ratio. The larger the prominence ratio, the more segmentations of a single entity category in the sentence, and the fewer segmentations of different entity categories in the sentence, and most of the different segmentations in the sentence are similar and repeated materials; the smaller the prominence ratio, the more evenly the number of segmentations of different entity categories in the sentence is distributed, ensuring the diversity of segmentations of different entity categories in the sentence.
[0059] The confusion ratio is recorded as R2, R2 = (P2 + L2 + O2) / S, the sum of P2, L2, and O2 is the entity confusion count. The larger the confusion ratio, the higher the frequency of confusing participles in the sentence, and there are multiple ways of understanding the sentence, which enhances the diversity of training data in subsequent model learning; the smaller the confusion ratio, the lower the frequency of confusing participles in the sentence, and the sentence only corresponds to a single way of understanding.
[0060] In one embodiment, clustering the semantic features of all the sentences in the corpus to be processed to determine redundant clusters includes: clustering the semantic features of all the sentences in the corpus to be processed to obtain multiple target clusters; determining the target cluster center coordinates of each target cluster based on all the semantic features in each target cluster; and determining the redundant cluster among the multiple target clusters based on the target cluster center coordinates of each target cluster.
[0061] Specifically, the semantic feature is composed of a prominence ratio and a confusion ratio, and is represented by a feature coordinate, denoted as (R1, R2). The semantic features of each sentence in the processed corpus are clustered, that is, each semantic feature is classified, and after classification, a target cluster composed of semantic features is obtained. Specifically, the clustering process can be implemented by the Kmeans algorithm, the K-center point algorithm or the CLARANS algorithm. In this embodiment, the Kmeans algorithm is selected to implement the clustering process.
[0062] The corresponding coordinates of all semantic features in the target cluster are averaged to determine the target cluster center coordinates of the target cluster, which are recorded as (X1, X2). That is, each target cluster corresponds to a different corresponding target cluster center coordinate. The center coordinates of each target cluster are arranged in ascending order according to (X1, -X2). The target cluster at the top of the sorting has the smallest prominence ratio and the largest confusion ratio. The sentences corresponding to the semantic features in this target cluster have word segmentations of multiple different entity categories, that is, this target cluster includes the data with the most learning and training value. After sorting in sequence, the target cluster at the end of the sorting has the largest prominence ratio and the smallest confusion ratio. The semantic features in this target cluster correspond to The frequency of confusing word segmentation in the sentence is low, which makes the understanding method of this type of sentence relatively simple and single, and the learning and training significance is low, that is, the target cluster includes the closest and redundant data, and the sentences corresponding to multiple semantic features in the target cluster are judged as repeated redundant corpus. The target cluster at the end of the sorting is taken as a redundant cluster, that is, all target clusters before the end of the sorting are retained, and the remaining target clusters except the first in the sorting can also be taken as redundant clusters, that is, the target clusters other than the most valuable for learning and training are regarded as redundant clusters. However, this method will lead to insufficient training data. The specific configuration can be made according to the training accuracy and training data amount of the subsequent model.
[0063] In one embodiment, the semantic features of all the sentences in the corpus to be processed are clustered to obtain multiple target clusters, including: obtaining initial cluster center coordinates of multiple initial clusters; forming feature coordinates with the prominence ratio and the confusion ratio of each sentence, and determining the Euclidean distance between each feature coordinate and each initial cluster center coordinate; among the multiple Euclidean distances corresponding to each feature coordinate, the initial cluster corresponding to the Euclidean distance with the smallest value is used as the target cluster of the feature coordinates.
[0064] Specifically, the number of clusters and the initial cluster center coordinates of each initial cluster are obtained. The number of clusters can be customized according to actual conditions. In this embodiment, the number of clusters is set to 4, that is, the four initial clusters are respectively recorded as C1, C2, C3, and C4, each initial cluster corresponds to an initial cluster center coordinate, and the feature coordinates are marked as (R1, R2). The Euclidean distance between the feature coordinates and the coordinates of each initial cluster center is calculated, and the initial cluster corresponding to the smallest Euclidean distance is used as the target cluster of the feature coordinates. The feature coordinates are divided into the target cluster, thereby completing the clustering processing of the feature coordinates corresponding to each sentence in the corpus to be processed.
[0065] In one embodiment, the obtaining of the corpus to be processed includes: obtaining original corpus; removing sentences in the original corpus that do not meet the length requirement and / or do not meet the character requirement to obtain the corpus to be processed.
[0066] Specifically, the original corpus is a corpus that has not been preprocessed, and the length requirement includes an upper length threshold and a lower length threshold. If the length of a sentence exceeds the length range from the lower length threshold to the upper length threshold, the sentence is determined to not meet the length requirement. The character requirement includes a preset text character. In this embodiment, the preset text character is a Chinese character. Sentences that do not contain Chinese characters are determined to be sentences that do not meet the character requirement. After removing the sentences that do not meet the length requirement and / or do not meet the character requirement from the original corpus, the preprocessed corpus to be processed is obtained.
[0067] Figure 1 FIG. 1 is a flow chart of a method for removing corpus redundancy in one embodiment. It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0068] In one embodiment, Figure 2 As shown, a corpus redundancy removal device is provided, comprising:
[0069] An acquisition module 210 is used to acquire a corpus to be processed, wherein the corpus to be processed includes a plurality of sentences, each of which includes a plurality of segmented words;
[0070] A recognition module 220, configured to determine the entity category of each of the word segments in each of the sentences;
[0071] A determination module 230, configured to determine the semantic features of each of the sentences based on the entity categories of all the segmented words in each of the sentences;
[0072] A clustering module 240 is used to perform clustering processing on the semantic features of all the sentences in the corpus to be processed to determine redundant clusters, wherein the semantic features in the redundant clusters are used to indicate the sentences with repeated semantics;
[0073] The removal module 250 is used to remove the sentences corresponding to each of the semantic features in the redundant cluster.
[0074] In one embodiment, the determination module 230 is specifically configured to:
[0075] According to the entity categories of all the participles in the sentence, counting entity confusion counts, the number of participles under each entity category, and the total number of participles under all entity categories, wherein the entity confusion count is used to indicate the frequency of the presence of a second participle under a second category in a first participle of a first category in the sentence, the first category and the second category are different entity categories, the first participle and the second participle are both participles in the sentence, and the length of the first participle is greater than the length of the second participle;
[0076] The semantic features of the sentence are determined according to the entity confusion count, the number of participles under each entity category, and the total number of participles under all entity categories.
[0077] In one embodiment, the determination module 230 is specifically configured to:
[0078] When there are multiple second participles of the second category in the first participle, the entity confusion count for the first participle is accumulated once.
[0079] In one embodiment, the determination module 230 is specifically configured to:
[0080] Among the number of participles under each entity category, the prominence ratio is obtained according to the ratio of the number of participles with the largest value to the total number of participles;
[0081] The confusion ratio is obtained according to the ratio of the entity confusion count to the total number of segmented words.
[0082] In one embodiment, the clustering module 240 is specifically used for:
[0083] Performing clustering processing on the semantic features of all the sentences in the corpus to be processed to obtain multiple target clusters;
[0084] Determining the target cluster center coordinates of each target cluster according to all the semantic features in each target cluster;
[0085] Among the multiple target clusters, the redundant cluster is determined according to the target cluster center coordinates of each target cluster.
[0086] In one embodiment, the clustering module 240 is specifically used for:
[0087] Obtain the initial cluster center coordinates of multiple initial clusters;
[0088] The prominence ratio and the confusion ratio of each of the sentences form feature coordinates, and determine the Euclidean distance between each of the feature coordinates and each of the initial cluster center coordinates;
[0089] Among the multiple Euclidean distances corresponding to the feature coordinates, the initial cluster corresponding to the Euclidean distance with the smallest value is used as the target cluster of the feature coordinates.
[0090] In one embodiment, the acquisition module 210 is specifically used for:
[0091] Get the original corpus;
[0092] The sentences that do not meet the length requirement and / or the character requirement are removed from the original corpus to obtain the corpus to be processed.
[0093] Figure 3 FIG. 1 shows an internal structure diagram of a computer device in an embodiment. The computer device may specifically be a server. Figure 3As shown, the computer device includes a processor, a memory, a network interface, an input device and a display screen connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement the method for removing corpus redundancy. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the method for removing corpus redundancy. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0094] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0095] In one embodiment, the corpus redundancy removal device provided in the present application can be implemented in the form of a computer program. Figure 3 The computer device is run on the computer device shown. The memory of the computer device can store various program modules that constitute the corpus redundancy removal device, such as: Figure 2 The acquisition module 210, identification module 220, determination module 230, clustering module 240 and removal module 250 are shown. The computer program composed of various program modules enables the processor to execute the steps of the corpus redundancy removal method of each embodiment of the present application described in this specification.
[0096] Figure 3 The computer device shown can be Figure 2The acquisition module 210 in the corpus redundancy removal device shown acquires the corpus to be processed, wherein the corpus to be processed includes a plurality of sentences, and each of the sentences includes a plurality of participles. The computer device may determine the entity category of each of the participles in each of the sentences through the identification module 220. The computer device may determine the semantic features of each of the sentences based on the entity categories of all the participles in each of the sentences through the determination module 230. The computer device may perform clustering processing on the semantic features of all the sentences in the corpus to be processed through the clustering module 240 to determine redundant clusters, wherein the semantic features in the redundant clusters are used to indicate the sentences with repeated semantics. The computer device may remove the sentences corresponding to each of the semantic features in the redundant clusters through the removal module 250.
[0097] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the method described in any one of the above embodiments is implemented when the processor executes the computer program.
[0098] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above embodiments is implemented.
[0099] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0100] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0101] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for removing redundancy from corpus, It is characterized in that The method comprises: Acquire a corpus to be processed, wherein the corpus to be processed includes a plurality of sentences, and each of the sentences includes a plurality of participles; Determining the entity category of each of the participles in each of the sentences, wherein the entity category is used to indicate the semantic category of the participle; Determining the semantic features of each of the sentences based on the entity categories of all the segmented words in each of the sentences, wherein the semantic features are represented by feature coordinates; The method comprises: performing clustering processing on the semantic features of all the sentences in the corpus to be processed to determine redundant clusters, including: performing clustering processing on the semantic features of all the sentences in the corpus to be processed to obtain multiple target clusters; determining the target cluster center coordinates of each target cluster according to all the semantic features in each target cluster; determining the redundant cluster according to the target cluster center coordinates of each target cluster among the multiple target clusters, wherein the semantic features in the redundant clusters are used to indicate the sentences with repeated semantics; The sentences corresponding to each of the semantic features in the redundant clusters are removed.
2. The method according to claim 1, It is characterized in that The determining of the semantic features of each of the sentences based on the entity categories of all the segmented words in each of the sentences includes: According to the entity categories of all the participles in the sentence, counting entity confusion counts, the number of participles under each entity category, and the total number of participles under all entity categories, wherein the entity confusion count is used to indicate the frequency of the presence of a second participle under a second category in a first participle of a first category in the sentence, the first category and the second category are different entity categories, the first participle and the second participle are both participles in the sentence, and the length of the first participle is greater than the length of the second participle; The semantic features of the sentence are determined according to the entity confusion count, the number of participles under each entity category, and the total number of participles under all entity categories.
3. The method according to claim 2, It is characterized in that The counting of entity confusion according to the entity categories of all the word segments in the sentence includes: When there are multiple second participles of the second category in the first participle, the entity confusion count for the first participle is accumulated once.
4. The method according to claim 3, It is characterized in that The semantic features include a prominence ratio and a confusion ratio, and determining the semantic features of the sentence according to the entity confusion count, the number of segmented words under each entity category, and the total number of segmented words under all entity categories includes: Among the number of participles under each entity category, the prominence ratio is obtained according to the ratio of the number of participles with the largest value to the total number of participles; The confusion ratio is obtained according to the ratio of the entity confusion count to the total number of segmented words.
5. The method according to claim 1, It is characterized in that The clustering process is performed on the semantic features of all the sentences in the corpus to be processed to obtain a plurality of target clusters, including: Obtain the initial cluster center coordinates of multiple initial clusters; The prominence ratio and confusion ratio of each of the sentences are used to form feature coordinates, and the Euclidean distance between each of the feature coordinates and each of the initial cluster center coordinates is determined; Among the multiple Euclidean distances corresponding to the feature coordinates, the initial cluster corresponding to the Euclidean distance with the smallest value is used as the target cluster of the feature coordinates.
6. The method according to claim 1, It is characterized in that The obtaining of the corpus to be processed includes: Get the original corpus; The sentences that do not meet the length requirement and / or the character requirement are removed from the original corpus to obtain the corpus to be processed.
7. A corpus redundancy removal device, It is characterized in that The device comprises: An acquisition module, used for acquiring a corpus to be processed, wherein the corpus to be processed includes a plurality of sentences, each of which includes a plurality of participles; A recognition module, used to determine the entity category of each of the segmented words in each of the sentences, wherein the entity category is used to indicate the semantic category of the segmented word; A determination module, configured to determine the semantic features of each of the sentences based on the entity categories of all the segmented words in each of the sentences, wherein the semantic features are represented by feature coordinates; A clustering module, used for clustering the semantic features of all the sentences in the corpus to be processed to determine redundant clusters, including: clustering the semantic features of all the sentences in the corpus to be processed to obtain multiple target clusters; determining the target cluster center coordinates of each target cluster according to all the semantic features in each target cluster; determining the redundant cluster among the multiple target clusters according to the target cluster center coordinates of each target cluster, wherein the semantic features in the redundant cluster are used to indicate the sentences with repeated semantics; A removal module is used to remove the sentences corresponding to each of the semantic features in the redundant cluster.
8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
News sentence clustering method based on semantic similarity, device and storage medium
CN107679144A