Corpus cleaning method and system, medium and terminal
Through pre-labeling, corpus optimization and enhanced processing, the problem of low corpus cleaning efficiency is solved, and the corpus quality and large model training efficiency are improved.
Patent Information
- Application Number
- CN202510758301.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-09
AI Technical Summary
In the prior art, the corpus cleaning efficiency is low and the corpus content is inconsistent after cleaning, which affects the training quality of the big model.
Through pre-labeling processing, corpus optimization processing and enhancement processing, multiple analytical information are generated, and combined with desensitization and standardization processing, the optimal target corpus is generated.
It improves the efficiency and quality of corpus cleaning and improves the training efficiency of large models.
Smart Images

Figure CN120278142A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data, and particularly relates to a corpus cleaning method, system, medium and terminal. Background Art
[0002] An artificial intelligence large model refers to a class of artificial intelligence models with a large number of parameters constructed by artificial neural networks. It usually first conducts pre-training on a large amount of data through self-supervised learning or semi-supervised learning, and then further optimizes its performance and capabilities through methods such as instruction tuning and alignment with humans. And the corpus is an important material for training artificial intelligence large models. Generally, the corpus refers to the instances and data sets used in linguistic research and natural language processing. It can be written texts, spoken records or other structured data, and is usually used for analyzing language phenomena, supporting machine translation, speech recognition, automatic text summarization and other tasks. Usually, the corpus is a set of text or speech data that has been collected, sorted and annotated. These data are used in linguistic research to analyze the usage rules of language, vocabulary changes and grammatical structures, etc. In natural language processing, the corpus is the basic data source for training and testing models, supporting functions such as machine translation, speech recognition, and sentiment analysis.
[0003] Due to the complex, non-standard and diverse characteristics of traditional corpora, before large model training, it is generally necessary to clean the corpus. Most traditional cleaning methods adopt a single processing method, resulting in problems such as non-compliance after corpus cleaning or inconsistent corpus content after corpus cleaning, affecting the quality of the corpus and being unfavorable for large model training. Summary of the Invention
[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a corpus cleaning method, system, medium and terminal, which is used to solve the problem of low efficiency of corpus cleaning in the prior art.
[0005] To achieve the above purpose and other related purposes, the present invention provides a corpus cleaning method, including the following steps: Obtain the original corpus to be cleaned, perform pre-annotation processing on the original corpus to generate a pre-annotated corpus, and parse the pre-annotated corpus to generate first parsing information; Perform corpus optimization on the original corpus to obtain an optimized corpus, and parse the optimized corpus to generate second parsing information; Perform enhancement processing on the original corpus to obtain an enhanced corpus, and parse the enhanced corpus to generate third parsing information; Generate target parsing information according to the first parsing information, the second parsing information and the third parsing information, and generate a corresponding synthetic corpus according to the target parsing information; Desensitize the synthetic corpus according to the desensitization library to generate a corresponding desensitized corpus, and generate a corresponding target corpus after standardizing the desensitized corpus.
[0006] In an embodiment of the present invention, perform pre-annotation processing on the original corpus to generate a pre-annotated corpus, and parse the pre-annotated corpus to generate first corpus information, including: Perform label grading on the original corpus to generate a tree structure, the tree structure includes a plurality of structural blocks, and each of the structural blocks corresponds to a field; Copy the original corpus multiple times to obtain multiple backup corpora, and perform hierarchical annotation on each backup corpus according to the tree structure according to a preset rule to obtain multiple backup annotation information; Perform differential comparison on multiple backup annotation information, and sequentially select the annotation word with the highest occurrence frequency at each position of the tree structure as the target annotation word, and combine multiple target annotation words together in order to form the pre-annotated corpus; Parse the pre-annotated corpus through a first parsing tool to obtain the first parsing information.
[0007] In an embodiment of the present invention, the differential comparison of multiple backup annotation information, sequentially select the annotation word with the highest occurrence frequency at each position of the tree structure as the target annotation word, and combine multiple target annotation words together in order to form the pre-annotated corpus, includes: Determine the structural block corresponding to each field in the original corpus according to the tree structure; At the position of each structural block, calculate the occurrence frequency of the annotation words that appear in the current structural block in each backup annotation information respectively, and select the annotation word corresponding to the maximum occurrence frequency as the target annotation word; Combine the target annotation words together in order according to the order of the structural blocks in the tree structure to form the pre-annotated corpus.
[0008] In an embodiment of the present invention, perform corpus optimization on the original corpus to obtain an optimized corpus, and parse the optimized corpus to generate second corpus information, including: Use the hash algorithm to calculate the data unique identifier of the original corpus and the semantic similarity in the original corpus, and remove duplicate content and approximate content according to the data unique identifier and the semantic similarity to obtain a deduplicated corpus; Perform syntax tree parsing on the deduplicated corpus to remove logical errors, and filter marketing content based on a rule template; After clearing the existing website characters, special characters and meaningless symbols, the optimized corpus is obtained; Parse the optimized corpus through a second parsing tool to obtain the second parsing information.
[0009] In an embodiment of the present invention, the enhancing the original corpus to obtain an enhanced corpus and parsing the enhanced corpus to generate third corpus information includes: Determine the non-core fields corresponding to the original corpus according to the tree structure; Generate a replacement corpus after replacing the non-core fields based on a domain vocabulary; Call a multi-language API combination to cyclically translate the replacement corpus to generate the enhanced corpus; Parse the enhanced corpus through a third parsing tool to obtain the third parsing information.
[0010] In an embodiment of the present invention, the generating target parsing information according to the first parsing information, the second parsing information, and the third parsing information, and generating a corresponding synthesized corpus according to the target parsing information includes: Split the first parsing information, the second parsing information, and the third parsing information according to the tree structure to obtain a first information segment, a second information segment, and a third information segment at corresponding positions; Calculate the information segment mean at each position according to the first information segment, the second information segment, and the third information segment; Calculate the coefficient of variation of the first information segment, the second information segment, and the third information segment at each position with respect to the information segment mean at the corresponding position, respectively; Select the one with the smallest coefficient of variation among the first information segment, the second information segment, and the third information segment as the target information segment; Arrange the target information segments at each position in order and combine them together to form the target parsing information, and generate the corresponding synthesized corpus according to the target parsing information.
[0011] In an embodiment of the present invention, the desensitizing the synthesized corpus according to a desensitization library to generate a corresponding desensitized corpus, and generating a corresponding target corpus after standardizing the desensitized corpus includes: Compare the synthesized corpus with the desensitization library to determine the sensitive words therein, and generate the desensitized corpus after replacing the sensitive words with encrypted words; Generate the corresponding target corpus after respectively performing format unification, lemmatization, and numerical standardization on the desensitized corpus.
[0012] The present invention also discloses a corpus cleaning system, including: The first parsing module is used to obtain the original corpus to be cleaned, perform pre-annotation processing on the original corpus to generate a pre-annotated corpus, and parse the pre-annotated corpus to generate first parsing information; The second parsing module is used to optimize the original corpus to obtain an optimized corpus, and parse the optimized corpus to generate second parsing information; The third parsing module is used to enhance the original corpus to obtain an enhanced corpus, and parse the enhanced corpus to generate third parsing information; The synthesis module is used to generate target parsing information according to the first parsing information, the second parsing information and the third parsing information, and generate a corresponding synthesized corpus according to the target parsing information; The processing module is used to desensitize the synthesized corpus according to the desensitization library to generate a corresponding desensitized corpus, and generate a corresponding target corpus after standardizing the decrypted corpus.
[0013] The present invention provides a storage medium on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned corpus cleaning method is implemented.
[0014] The present invention provides a terminal, including: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program stored in the memory, so that the terminal executes the above-mentioned corpus cleaning method.
[0015] As described above, the corpus cleaning method, system, medium and terminal of the present invention have the following beneficial effects: The present invention pre-annotates, optimizes and enhances the original corpus to be cleaned respectively and parses them to obtain the first parsing information, the second parsing information and the third parsing information, cleans and optimizes the corpus from different directions, and then generates target parsing information according to the first parsing information, the second parsing information and the third parsing information, generates a processed synthesized corpus according to the target parsing information, and finally obtains the target corpus through desensitization and standardization processing, thereby optimizing the original corpus simultaneously from three levels of corpus annotation, corpus optimization and corpus enhancement, and finally selecting the optimal target corpus, thus effectively realizing the cleaning of the original corpus, ensuring the corpus cleaning efficiency while improving the quality of the original corpus and also improving the training efficiency of the large model. Description of the Drawings
[0016] Figure 1 It shows a flowchart of the corpus cleaning of the present invention in an embodiment.
[0017] Figure 2 It shows a specific process diagram of step S400 in the corpus cleaning of the present invention.
[0018] Figure 3 Shown is a structural block diagram of the corpus cleaning system of the present invention in an embodiment. Detailed implementation manners
[0019] The following uses specific specific embodiments to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0020] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. The diagrams only show the components related to the present invention and are not drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0021] The corpus cleaning method, system, medium, and terminal of the present invention perform pre-annotation processing, corpus optimization processing, and enhancement processing on the original corpus to be cleaned respectively and then parse to obtain the first parsing information, the second parsing information, and the third parsing information, clean and optimize the corpus from different directions, and then generate target parsing information according to the first parsing information, the second parsing information, and the third parsing information, generate the processed synthetic corpus according to the target parsing information, and finally obtain the target corpus through desensitization and standardization processing, thereby optimizing the original corpus from three levels of corpus annotation, corpus optimization, and corpus enhancement, and finally selecting the optimal target corpus, thus effectively realizing the cleaning of the original corpus, ensuring the corpus cleaning efficiency while improving the quality of the original corpus and also improving the training efficiency of the large model.
[0022] A computer program is stored on the storage medium of the present invention, and when the computer program is executed by a processor, the following corpus cleaning method is implemented. The storage medium includes: various media that can store program codes such as read-only memory (ROM), random access memory (RAM), magnetic disk, USB flash drive, memory card, or optical disc.
[0023] Any combination of one or more storage media may be employed. The storage media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present document, a computer-readable storage medium may be any tangible medium that contains or stores a program which can be used by or in connection with an instruction execution system, apparatus, or device.
[0024] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take many forms, including - but not limited to - an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium that is not a computer-readable storage medium and that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0025] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including - but not limited to - wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0026] The computer program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages - such as Java, Smalltalk, C++ - and conventional procedural programming languages - such as the "C" programming language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0027] The present invention will now be described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems) and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, thereby producing a machine such that when the computer program instructions are executed by the processor of the computer or other programmable data processing apparatus, an apparatus is created that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0028] These computer program instructions can also be stored in a computer-readable medium, which can direct a computer, other programmable data processing apparatus, or other devices to operate in a particular manner, so that the instructions stored in the computer-readable medium produce an article of manufacture including instructions that implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0029] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices, such that a series of operation steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide a process that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0030] The terminal of the present invention includes a processor and a memory.
[0031] The memory is used to store computer programs; preferably, the memory includes various media such as ROM, RAM, magnetic disks, USB flash drives, memory cards, or optical discs that can store program codes.
[0032] The processor is connected to the memory and is used to execute the computer programs stored in the memory, so that the terminal executes the following corpus cleaning method.
[0033] Preferably, the processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0034] As Figure 1 shown, in one embodiment, the present invention discloses a corpus cleaning method, including the following steps: S100. Obtain the original corpus to be cleaned, perform pre-annotation processing on the original corpus to generate a pre-annotated corpus, and parse the pre-annotated corpus to generate first parsing information.
[0035] In some embodiments, performing pre-annotation processing on the original corpus to generate a pre-annotated corpus and parsing the pre-annotated corpus to generate first corpus information includes: Perform label grading on the original corpus to generate a tree structure, the tree structure includes a plurality of structural blocks, and each structural block corresponds to a field; Copy the original corpus multiple times to obtain multiple copies of the corpus, and according to a preset rule, perform hierarchical annotation on each copy of the corpus according to the tree structure to obtain multiple backup annotation information; Perform differential comparison on multiple pieces of the backup annotation information, and sequentially select the annotation word with the highest frequency of occurrence at each position of the tree structure as the target annotation word, and combine multiple target annotation words together in sequence to form the pre-annotated corpus; Parse the pre-annotated corpus through a first parsing tool to obtain the first parsing information.
[0036] In this embodiment, first obtain the original corpus to be cleaned, and then perform pre-annotation processing on the original corpus to obtain a pre-annotated corpus, so as to facilitate parsing the pre-annotated corpus to generate first parsing information, which is convenient for performing synthetic cleaning on the original corpus according to the first parsing information in the subsequent step to improve the quality of the original corpus.
[0037] Specifically, first, the original corpus is hierarchically tagged according to the structure of the original corpus to generate a tree structure corresponding to the original corpus. The tree structure includes multiple structural blocks, each structural block corresponding to a field. The fields include core fields and non-core fields. The structural block corresponding to the core field serves as the main part of the tree structure, and the structural block corresponding to the non-core field serves as the branch part of the structural block corresponding to the core field.
[0038] After that, multiple copies of the original corpus are obtained to get multiple backup corpora, and according to the preset annotation rules, each backup corpus is hierarchically annotated according to the tree structure to obtain corresponding multiple backup annotation information. Then, the obtained multiple backup annotation information is differentially compared, and the annotation word with the highest frequency of occurrence is selected as the target annotation word at each position of the tree structure in turn, and the target annotation words are combined together in the order of the tree structure to form the pre-annotated corpus, and the pre-annotated corpus is parsed by the first parsing tool to obtain the first parsing information, which is convenient for subsequent cleaning of the original corpus according to the first parsing information.
[0039] Among them, before hierarchically annotating each backup corpus according to the tree structure, it also includes parsing the logical errors existing in the backup corpus according to the syntax tree and filtering out auxiliary words and interjections to ensure the accuracy of subsequent annotation results.
[0040] In the above process, by copying the original corpus to obtain multiple backup corpora, then hierarchically annotating each backup corpus according to the tree structure to obtain multiple backup annotation information, differentially comparing the multiple backup annotation information in the subsequent process, and sequentially selecting the annotation word with the highest frequency of occurrence as the target annotation word, and then combining the target annotation words at each position together to form the final pre-annotated corpus. Compared with the traditional annotation method, the solution of this application ensures the accuracy of the pre-annotated corpus by performing multiple annotation processes on the original corpus and selecting the optimal annotation words to form a complete pre-annotated corpus, which is beneficial to improving the quality of corpus cleaning.
[0041] In some other embodiments, the differentially comparing the multiple backup annotation information, and sequentially selecting the annotation word with the highest frequency of occurrence as the target annotation word at each position of the tree structure, and combining the multiple target annotation words together in order to form the pre-annotated corpus includes: Determining the structural block corresponding to each field in the original corpus according to the tree structure; At the position of each structural block, calculate the frequency of occurrence of the annotation words that appear in the current structural block in each of the backup annotation information, and select the annotation word corresponding to the maximum frequency of occurrence as the target annotation word; Combining the target annotation words together in sequence according to the order of the structural blocks in the tree structure to form the pre-annotated corpus.
[0042] In this embodiment, after hierarchical annotation of multiple backup corpora to obtain multiple backup annotation information, first determine the structural block corresponding to each field in the original corpus according to the tree structure, and at the position of each structural block, calculate the occurrence frequency of the annotation words of the multiple backup annotation information in the current structural block respectively, and select the annotation word with the largest occurrence frequency as the target annotation word corresponding to the current structural block. After calculating and selecting one by one according to the order of the tree structure, the target annotation word corresponding to each structural block is obtained. Then, according to the order of the structural blocks in the tree structure, the target annotation words corresponding to the structural blocks are combined together in sequence to form the pre-annotated corpus.
[0043] Wherein, the occurrence frequency is the ratio of the number of occurrences of the target word at the same position in the tree structure to the total number of the backup corpora.
[0044] S200. Optimize the original corpus to obtain an optimized corpus, and parse the optimized corpus to generate second parsing information.
[0045] In some embodiments, the optimizing the original corpus to obtain an optimized corpus and parsing the optimized corpus to generate second corpus information includes: Calculating the data unique identifier of the original corpus and the semantic similarity in the original corpus by using a hash algorithm, and removing duplicate content and approximate content according to the data unique identifier and the semantic similarity to obtain a deduplicated corpus; Performing syntactic tree parsing on the deduplicated corpus to remove logical errors, and then filtering marketing content based on a rule template; After clearing the existing website characters, special characters and meaningless symbols, the optimized corpus is obtained; Parsing the optimized corpus through a second parsing tool to obtain the second parsing information.
[0046] In this embodiment, in order to optimize the original corpus, first, a unique data identifier corresponding to each character in the original corpus is calculated through a hash algorithm, and the semantic similarity of the characters in the original corpus is calculated according to the MinHash algorithm or the SimHash algorithm. Then, duplicate content is removed based on the data unique identifier, and approximate content is removed based on the semantic similarity to obtain the deduplicated corpus. After that, the deduplicated corpus is parsed through a syntax tree analysis to remove logical errors, and then the marketing content of the deduplicated corpus is filtered based on a rule template. After removing the website characters, special characters, and meaningless symbols existing in the deduplicated corpus, the final optimized corpus is obtained. Then, the optimized corpus is semantically parsed through a second parsing tool to obtain second parsing information.
[0047] S300. Perform enhancement processing on the original corpus to obtain an enhanced corpus, and parse the enhanced corpus to generate third parsing information.
[0048] In some other embodiments, the performing enhancement processing on the original corpus to obtain an enhanced corpus and parsing the enhanced corpus to generate third corpus information includes: Determine the non-core fields corresponding to the original corpus according to the tree structure; Generate a replacement corpus after replacing the non-core fields based on a domain word list; Call a multi-language API combination to cyclically translate the replacement corpus to generate the enhanced corpus; Parse the enhanced corpus through a third parsing tool to obtain the third parsing information.
[0049] In this embodiment, first, the non-core fields corresponding to the original corpus are determined according to the tree structure, that is, the fields corresponding to the structural blocks of the branch parts corresponding to the main trunk part in the tree structure. A replacement corpus is generated after replacing the non-core fields based on the domain word list, so as to ensure that the original corpus is replaced while retaining the original semantics. At the same time, a multi-language API combination is called to cyclically translate the replacement corpus to generate the enhanced corpus. For example, if the original corpus is in Chinese, the multi-language API is called to translate the replacement corpus into English, then into Japanese, and then back into the original Chinese to generate a new sample with the same semantics as the original corpus, that is, the enhanced corpus. Then, based on the obtained enhanced corpus, a third parsing tool is called to parse the enhanced corpus to obtain third parsing information.
[0050] It should be noted that in the above-mentioned respective parsing processes, the parsing processes of the first parsing tool, the second parsing tool, and the third parsing tool are the content of the prior art, mainly to obtain the semantic information of the corpus through parsing. The selection of the first parsing tool, the second parsing tool, and the third parsing tool can also be selected according to actual needs. This solution does not limit the tool itself and will not be elaborated here.
[0051] S400. Generate target parsing information according to the first parsing information, the second parsing information, and the third parsing information, and generate a corresponding synthetic corpus according to the target parsing information.
[0052] In some embodiments, referring to Figure 2 , the generating target parsing information according to the first parsing information, the second parsing information, and the third parsing information, and generating a corresponding synthetic corpus according to the target parsing information includes: S401. Split the first parsing information, the second parsing information, and the third parsing information according to the tree structure to obtain the first information segment, the second information segment, and the third information segment at the corresponding positions; S402. Calculate the information segment mean value at each position according to the first information segment, the second information segment, and the third information segment; S403. Calculate the coefficient of variation of the first information segment, the second information segment, and the third information segment at each position from the information segment mean value at the corresponding position respectively; S404. Select the one with the smallest coefficient of variation among the first information segment, the second information segment, and the third information segment as the target information segment; S405. Arrange and combine the target information segments at each position in order to form the target parsing information, and generate the corresponding synthetic corpus according to the target parsing information.
[0053] In this embodiment, after obtaining the first parsing information, the second parsing information, and the third parsing information corresponding to the original corpus respectively, according to the tree structure of the original corpus, the first parsing information, the second parsing information, and the third parsing information are independently split respectively to obtain a plurality of first information segments, second information segments, and third information segments with the same number as the number of the structural blocks. Then, according to the first information segment, the second information segment, and the third information segment, the information segment mean corresponding to the position of each structural block is calculated, and the coefficient of variation of the first information segment, the second information segment, and the third information segment at the corresponding position of each structural block from the information segment mean at the corresponding position is calculated respectively. Then, at the positions of each of the structural blocks, the information segment corresponding to the minimum coefficient of variation is selected as the target information segment, so as to obtain target information segments with the same number as the number of the structural blocks. And according to the order of the structural blocks in the tree structure, a plurality of target information segments are arranged and combined in sequence to form the target parsing information, and a corresponding synthetic corpus is generated according to the target parsing information, so as to realize the conversion processing of the corpus according to the parsing information, ensuring the quality of the finally generated synthetic corpus.
[0054] It should be noted that in the above process of calculating the coefficient of variation, for the convenience of calculation, the first information segment, the second information segment, and the third information segment corresponding to each structural block are all converted into corresponding coding values, such as ASCII coding, GB2312, GBK, Unicode coding. When calculating the coefficient of variation, the information segment mean corresponding to the current structural block is calculated through the coding values corresponding to the first information segment, the second information segment, and the third information segment respectively, and the coefficient of variation can be obtained by calculating the ratio of the absolute value of the difference between the coding values corresponding to the first information segment, the second information segment, and the third information segment and the information segment mean to the information segment mean.
[0055] Exemplarily, the coefficient of variation satisfies the following calculation formula: , where i is an integer from 1 to 3.
[0056] Where, represents the coefficient of variation, represents the information segment mean, represents the coding value corresponding to the first information segment, the coding value corresponding to the second information segment, the coding value corresponding to the third information segment.
[0057] S500. Desensitize the synthetic corpus according to the desensitization library to generate a corresponding desensitized corpus, and generate a corresponding target corpus after standardizing the desensitized corpus.
[0058] In some embodiments, desensitizing the synthetic corpus according to the desensitization library to generate a corresponding desensitized corpus, and generating a corresponding target corpus after standardizing the desensitized corpus, includes: Comparing the synthetic corpus with the desensitization library to determine sensitive words therein, and generating the desensitized corpus after replacing the sensitive words with encrypted words; After uniformly formatting, lemmatizing, and numerically standardizing the desensitized corpus, generating the corresponding target corpus.
[0059] In an embodiment, after obtaining the synthetic corpus, desensitize the synthetic corpus according to the desensitization library, generate a desensitized corpus after replacing sensitive words with encrypted words, and after uniformly formatting, magnetically normalizing, and numerically standardizing the desensitized corpus, generate a corresponding target corpus, thereby completing the cleaning work of the corpus.
[0060] Among them, lemmatization mainly targets the text data in the synthetic corpus, and numerical standardization targets the numerical data in the synthetic corpus. Since lemmatization and numerical standardization are the content of the prior art, this solution does not make special limitations on them and will not be elaborated here.
[0061] It should be noted that the protection scope of the method described in the present invention is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or subtracting steps of the prior art and replacing steps according to the principle of the present invention is included in the protection scope of the present invention.
[0062] The present invention also provides a corpus cleaning system. Referring to FIG. 3, it includes: A first parsing module 301, configured to obtain the original corpus to be cleaned, perform pre-annotation processing on the original corpus to generate a pre-annotated corpus, and parse the pre-annotated corpus to generate first parsing information; A second parsing module 302, configured to optimize the original corpus to obtain an optimized corpus, and parse the optimized corpus to generate second parsing information; A third parsing module 303, configured to enhance the original corpus to obtain an enhanced corpus, and parse the enhanced corpus to generate third parsing information; A synthesis module 304, configured to generate target parsing information according to the first parsing information, the second parsing information, and the third parsing information, and generate a corresponding synthetic corpus according to the target parsing information; A processing module 305, configured to desensitize the synthetic corpus according to the desensitization library to generate a corresponding desensitized corpus, and generate a corresponding target corpus after standardizing the decrypted corpus.
[0063] Among them, since the structure and principle of the corpus cleaning system correspond one by one to the steps in the above-mentioned corpus cleaning method, they will not be elaborated here.
[0064] It should be noted that it should be understood that the division of each module of the above system is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the x module can be a separately established processing element, or can be integrated in a certain chip of the above system. In addition, it can also be stored in the memory of the above system in the form of program code, and called and executed by a certain processing element of the above system to perform the functions of the above x module. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit of the hardware in the processor element or the instructions in the form of software.
[0065] For example, the above-mentioned modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-On-a-Chip (SOC).
[0066] It should be noted that the corpus cleaning system of the present invention can implement the method of the present invention, but the implementation device of the corpus cleaning method of the present invention includes but is not limited to the structure of the corpus cleaning system listed in this embodiment. Any structural deformation and substitution of the prior art made according to the principle of the present invention are included in the protection scope of the present invention.
[0067] The present invention discloses a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned corpus cleaning method is implemented.
[0068] The present invention also discloses a terminal, including: a processor and a memory; The memory is used for storing a computer program; The processor is used for executing the computer program stored in the memory, so that the terminal executes the above-mentioned corpus cleaning method.
[0069] In summary, for the corpus cleaning method, system, medium and terminal of the present invention, the original corpus to be cleaned is respectively subjected to pre-annotation processing, corpus optimization processing and enhancement processing and then parsed to obtain first parsing information, second parsing information and third parsing information, and the corpus is cleaned and optimized from different directions. Subsequently, target parsing information is generated according to the first parsing information, the second parsing information and the third parsing information, and a processed synthetic corpus is generated according to the target parsing information, and finally the target corpus is obtained through desensitization and standardization processing. Thus, the original corpus is optimized from three aspects of corpus annotation, corpus optimization and corpus enhancement, and finally the optimal target corpus is selected, effectively realizing the cleaning of the original corpus, ensuring the corpus cleaning efficiency while improving the quality of the original corpus and also improving the training efficiency of the large model. Therefore, the present invention effectively overcomes various drawbacks in the prior art and has high industrial utilization value.
[0070] The above embodiments are only illustrative of the principles and effects of the present invention, and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A corpus cleaning method, characterized in that, It includes the following steps: Obtain the original corpus to be cleaned, perform pre-annotation processing on the original corpus to generate a pre-annotated corpus, and parse the pre-annotated corpus to generate first parsing information; Optimize the original corpus to obtain an optimized corpus, and parse the optimized corpus to generate second parsing information; Perform enhancement processing on the original corpus to obtain an enhanced corpus, and parse the enhanced corpus to generate third parsing information; Generate target parsing information based on the first parsing information, the second parsing information, and the third parsing information, and generate a corresponding synthetic corpus based on the target parsing information; Perform desensitization processing on the synthetic corpus according to a desensitization library to generate a corresponding desensitized corpus, and generate a corresponding target corpus after standardizing the desensitized corpus.
2. The corpus cleaning method according to claim 1, characterized in that, Perform pre-annotation processing on the original corpus to generate a pre-annotated corpus, and parse the pre-annotated corpus to generate first corpus information, including: Perform label grading on the original corpus to generate a tree structure, the tree structure includes multiple structural blocks, and each structural block corresponds to a field; Copy the original corpus multiple times to obtain multiple backup corpora, and perform hierarchical annotation on each backup corpus according to a preset rule based on the tree structure to obtain multiple backup annotation information; Perform differential comparison on the multiple backup annotation information, and sequentially select the annotation word with the highest occurrence frequency at each position of the tree structure as the target annotation word, and combine the multiple target annotation words together in order to form the pre-annotated corpus; Parse the pre-annotated corpus through a first parsing tool to obtain the first parsing information.
3. The corpus cleaning method according to claim 2, wherein The performing differential comparison on the multiple backup annotation information, and sequentially selecting the annotation word with the highest occurrence frequency at each position of the tree structure as the target annotation word, and combining the multiple target annotation words together in order to form the pre-annotated corpus includes: Determine the structural block corresponding to each field in the original corpus according to the tree structure; At the position of each structural block, calculate the occurrence frequency of the annotation words that appear in the current structural block in each of the backup annotation information, and select the annotation word corresponding to the maximum occurrence frequency as the target annotation word; Combine the target annotation words together in order according to the order of the structural blocks in the tree structure to form the pre-annotated corpus.
4. The corpus cleaning method according to claim 1, wherein The optimizing the original corpus to obtain an optimized corpus, and parsing the optimized corpus to generate second corpus information includes: Calculate the data unique identifier of the original corpus and the semantic similarity in the original corpus by using a hash algorithm, and remove duplicate content and approximate content according to the data unique identifier and the semantic similarity to obtain a deduplicated corpus; Perform syntax tree parsing on the deduplicated corpus to remove logical errors, and filter marketing content based on a rule template; Remove the existing website characters, special characters, and meaningless symbols to obtain the optimized corpus; Parse the optimized corpus through a second parsing tool to obtain the second parsing information.
5. The corpus cleaning method according to claim 1, wherein Performing enhancement processing on the original corpus to obtain an enhanced corpus, and parsing the enhanced corpus to generate third corpus information, including: Determining non-core fields corresponding to the original corpus according to the tree structure; Generating a replacement corpus after replacing the non-core fields based on a domain vocabulary; Invoking a multilingual API combination to cyclically translate the replacement corpus to generate the enhanced corpus; Parsing the enhanced corpus through a third parsing tool to obtain the third parsing information.
6. The corpus cleaning method according to claim 2, wherein Generating target parsing information according to the first parsing information, the second parsing information, and the third parsing information, and generating a corresponding synthetic corpus according to the target parsing information, including: Splitting the first parsing information, the second parsing information, and the third parsing information according to the tree structure to obtain first information segments, second information segments, and third information segments at corresponding positions; Calculating the information segment mean values at each position according to the first information segment, the second information segment, and the third information segment; Calculating the coefficient of variation of the first information segment, the second information segment, and the third information segment at each position with respect to the information segment mean value at the corresponding position respectively; Selecting the one with the smallest coefficient of variation among the first information segment, the second information segment, and the third information segment as the target information segment; Arranging and combining the target information segments at each position in order to form the target parsing information, and generating the corresponding synthetic corpus according to the target parsing information.
7. The corpus cleaning method according to claim 1, wherein Performing desensitization processing on the synthetic corpus according to a desensitization library to generate a corresponding desensitized corpus, and performing standardization processing on the desensitized corpus to generate a corresponding target corpus, including: Comparing the synthetic corpus with the desensitization library to determine sensitive vocabulary therein, and generating the desensitized corpus after replacing the sensitive vocabulary with encrypted vocabulary; Performing format unification, word form normalization processing, and numerical standardization processing on the desensitized corpus respectively to generate the corresponding target corpus.
8. A corpus cleaning system, characterized in that, Including: A first parsing module, configured to obtain an original corpus to be cleaned, perform pre-annotation processing on the original corpus to generate a pre-annotated corpus, and parse the pre-annotated corpus to generate first parsing information; A second parsing module, configured to perform corpus optimization on the original corpus to obtain an optimized corpus, and parse the optimized corpus to generate second parsing information; A third parsing module, configured to perform enhancement processing on the original corpus to obtain an enhanced corpus, and parse the enhanced corpus to generate third parsing information; A synthesis module, configured to generate target parsing information according to the first parsing information, the second parsing information, and the third parsing information, and generate a corresponding synthetic corpus according to the target parsing information; A processing module, configured to perform desensitization processing on the synthetic corpus according to a desensitization library to generate a corresponding desensitized corpus, and perform standardization processing on the decrypted corpus to generate a corresponding target corpus.
9. A storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the corpus cleaning method according to any one of claims 1 to 7.
10. A terminal, characterized in that, Including: A processor and a memory; The memory is used to store a computer program; The processor is used to execute the computer program stored in the memory, so that the terminal executes the corpus cleaning method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Corpus cleaning method, device and equipment and medium
CN109739956A
Text processing method and device, equipment and storage medium
CN113139033A
Language model training method and device for device tree clustering solution
CN117332132A
Quality improvement method and equipment based on corpus cleaning, medium and product
CN119203996A
Corpus management method and system
CN119719254A