Corpus management method and system

Through a systematic corpus management process, including selection of collection channels, preprocessing, cleaning, screening, labeling and encapsulating data packets, and maintaining the encapsulated data packets, the problem of irregular corpus management in the existing technology is solved, and the corpus quality and user experience are improved.

CN119719254BActive Publication Date: 2025-06-06SHANGHAI INNOVATION INSTITUTE FOR SMART PROCESS MANUFACTURING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510222525.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-06
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

The lack of systematic management of the entire process of the corpus in the prior art has led to irregular use of corpus and affecting user experience.

Method used

Provide a corpus management method, including selecting a collection channel according to corpus requirements to obtain the original corpus, performing pre-processing, corpus cleaning, alignment screening, labeling and encapsulating data packets, and maintaining the encapsulated data packets.

Benefits of technology

Through a systematic corpus management process, improve corpus quality and user experience, and ensure the safety and long-term service life of corpus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719254B_ABST
    Figure CN119719254B_ABST
Patent Text Reader

Abstract

The present invention provides a corpus management method and system, the method comprising: according to the corpus requirements, selecting a corresponding collection channel to obtain the original corpus, preprocessing the original corpus to obtain an intermediate corpus; according to a preset standard, corpus cleaning the intermediate corpus to obtain a cleaned corpus; according to a value alignment rule, aligning and screening the cleaned corpus to obtain a target corpus; corpus annotation of the target corpus to obtain annotation data corresponding to each of the target corpuses; according to the format of the target corpus, encapsulating the target corpus and the annotation data together to obtain an encapsulated data packet, and maintaining the encapsulated data packet. The present invention effectively manages the corpus, can improve the use efficiency of the corpus, and enhance the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of big data, and in particular relates to a corpus management method and system. Background Art

[0002] Artificial intelligence big models refer to "large parameter" models trained using large-scale data and powerful computing power. These models are usually highly versatile and generalized, and can be applied to natural language processing, image recognition, speech recognition and other fields. They can be divided into large language models, large visual models, multimodal large models, and basic large models.

[0003] Corpus is an important material used to train large artificial intelligence models. Corpus generally refers to examples and data sets used in linguistic research and natural language processing. It can be written text, oral records or other structured data, usually used to analyze language phenomena, support machine translation, speech recognition, automatic text summarization and other tasks. A common corpus is a set of text or speech data that has been collected, organized and annotated. These data are used in linguistic research to analyze language usage patterns, vocabulary changes and grammatical structures. In natural language processing, corpus is the basic data source for training and testing models, supporting functions such as machine translation, speech recognition, and sentiment analysis.

[0004] The existing technology lacks systematic management of the entire corpus process, which easily leads to the problem of irregular use of corpus. Moreover, the corpus is often managed in a fixed way, which cannot be adjusted in time according to the use of the corpus, which has a negative impact on the user experience. Summary of the invention

[0005] In view of the above-mentioned shortcomings of the prior art, an object of the present invention is to provide a corpus management method and system for solving the problems of irregular corpus management and user experience impact in the prior art.

[0006] To achieve the above objectives and other related objectives, the present invention provides a corpus management method, comprising:

[0007] According to the corpus requirements, select the corresponding collection channel to obtain the original corpus, and preprocess the original corpus to obtain the intermediate corpus;

[0008] According to a preset standard, the intermediate corpus is cleaned to obtain a cleaned corpus;

[0009] According to the value alignment rule, the cleaned corpus is aligned and screened to obtain the target corpus;

[0010] Performing corpus annotation on the target corpus to obtain annotation data corresponding to each target corpus;

[0011] According to the format of the target corpus, the target corpus and the annotated data are encapsulated together to obtain an encapsulated data packet, and the encapsulated data packet is maintained.

[0012] In one embodiment of the present invention, the collection channel includes a public data source channel, a private procurement channel and a shared collection channel, and the step of selecting a corresponding collection channel to obtain the original corpus according to the corpus requirements includes:

[0013] Obtaining the user's corpus requirement, and determining the type of the required original corpus according to the corpus requirement;

[0014] Selecting a corresponding collection channel according to the source type of the original corpus to obtain the original corpus;

[0015] Wherein, when the source type of the original corpus is a public type, selecting a public data source channel to obtain the original corpus;

[0016] When the source type of the original corpus is a private type, the private procurement channel is selected to obtain the original corpus; when the source type of the original corpus is a shared type, the shared acquisition channel is selected to obtain the original corpus.

[0017] In one embodiment of the present invention, the preprocessing of the original corpus to obtain the intermediate corpus includes:

[0018] Evenly split the original corpus into multiple corpus groups, and perform quality check on each of the corpus groups;

[0019] Obtain data integrity, data consistency, data accuracy and data timeliness of each group of the corpus for evaluation based on the quality check results;

[0020] Calculate a preliminary evaluation value for each of the corpus groups according to the evaluation results of the data completeness, the data consistency, the data accuracy and the data timeliness;

[0021] The corpus groups whose preliminary evaluation values ​​are less than a first threshold are eliminated, and the remaining corpus groups are recombined together to obtain the intermediate corpus.

[0022] In one embodiment of the present invention, the step of performing corpus cleaning on the intermediate corpus according to a preset standard to obtain a cleaned corpus includes:

[0023] Determine the target format according to the preset standard, and convert the intermediate corpus into the target format;

[0024] Performing data cleaning on the intermediate corpus after the format conversion, and performing normalization adjustment after the data cleaning to obtain the first corpus;

[0025] After completing the normalization adjustment, performing balancing processing on the first corpus to obtain a second corpus;

[0026] The second corpus is segmented and deduplicated to obtain cleaned corpuses, wherein each of the cleaned corpuses includes multiple data sets.

[0027] In one embodiment of the present invention, the step of aligning and screening the cleaned corpus according to the value alignment rule to obtain the target corpus includes:

[0028] The cleaned corpus is subjected to ethical screening, legal screening and commercial screening in turn, and the corpus that does not meet the ethical requirements, legal requirements and commercial requirements is removed according to the screening results to obtain the target corpus.

[0029] In one embodiment of the present invention, the tagging of the target corpus to obtain the tagging data corresponding to each target corpus includes:

[0030] Classifying the target corpus according to the data type of the target corpus to obtain classified corpus, and formulating corresponding annotation specifications for the classified corpus;

[0031] Performing data preprocessing on each of the classified corpora according to the annotation specification and then segmenting the data to obtain multiple data sets;

[0032] Preliminarily annotating each of the data sets to obtain primary annotated data;

[0033] The primary annotated data are reviewed and feedback is provided, the primary annotated data are modified and adjusted according to the feedback, and the adjusted primary annotated data are combined together to obtain the annotated data corresponding to the target corpus.

[0034] In one embodiment of the present invention, encapsulating the target corpus and the annotated data together to obtain an encapsulated data packet according to the format of the target corpus includes:

[0035] After arranging and backing up the classified corpora, determining a packaging standard and packaging format for each of the classified corpora;

[0036] Encapsulating the classified corpus and the annotated data according to the encapsulation standard and the encapsulation format to obtain encapsulated data;

[0037] According to the set metadata structure, metadata corresponding to the packaged data is generated during the packaging process, and the metadata is used to record relevant information of the packaged data, including data source, collection time, annotation information and usage authority;

[0038] The encapsulated data and the metadata are combined together and compressed and encrypted to form an encapsulated data packet.

[0039] In one embodiment of the present invention, in the process of generating the encapsulated data packet, the method further includes:

[0040] Writing a packaging description document, recording relevant parameters of the packaging data packet through the packaging description document;

[0041] Generate an encapsulation log to record the operation content of the encapsulated data packet through the encapsulation log;

[0042] Among them, the relevant parameters include the directory structure, file naming rules, compression and encryption methods, encapsulation standards and formats, metadata definitions and encapsulation processes of the encapsulated data, and the encapsulation log includes encapsulation steps, operation logs, operation time, operators and operation content.

[0043] In one embodiment of the present invention, the maintaining of the encapsulated data packet includes:

[0044] Acquire usage information and feedback information of the target corpus in each of the encapsulated data packets, and calculate the value parameter of the target corpus according to the usage information and the feedback information;

[0045] Regularly adjusting the order of the target corpus in the encapsulated data packet according to the value parameter, and regenerating a new encapsulated data packet according to the adjusted target corpus;

[0046] Adjusting the storage period of the target corpus according to the value parameter;

[0047] The greater the value parameter of the target corpus is, the longer the corresponding adjusted storage period is.

[0048] The present invention also provides a corpus management system, comprising:

[0049] A preprocessing module is used to select a corresponding collection channel to obtain original corpus according to corpus requirements, and preprocess the original corpus to obtain intermediate corpus;

[0050] A cleaning module, used for performing corpus cleaning on the intermediate corpus according to a preset standard to obtain a cleaned corpus;

[0051] A screening module, used for performing alignment screening on the cleaned corpus according to a value alignment rule to obtain a target corpus;

[0052] A tagging module, used for tagging the target corpus to obtain tagging data corresponding to each target corpus;

[0053] The encapsulation maintenance module is used to encapsulate the target corpus and the annotated data together according to the format of the target corpus to obtain an encapsulated data packet, and maintain the encapsulated data packet.

[0054] As described above, the corpus management method and system of the present invention have the following beneficial effects:

[0055] The present invention can select corresponding collection channels to obtain original corpora according to corpus requirements, and after obtaining the original corpora, pre-process the original corpora to improve subsequent processing efficiency; in the subsequent processing process, the original corpora are screened again by corpus cleaning and screening to obtain target corpora to improve corpus quality, and the target corpora are marked to obtain marked data, which is convenient for subsequent users to understand when using the target corpora, and improves the efficiency and convenience of users using the target corpora, and according to different formats of the target corpora, different target corpora and corresponding marked data are packaged to obtain packaged data packets, which is convenient for users to call and ensures the security of the corpora; and the packaged data packets are maintained throughout the entire use cycle, which can effectively extend the service life of the corpora. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 Shown is a flow chart of the corpus management method of the present invention.

[0057] Figure 2 Shown is a structural block diagram of the corpus management system of the present invention. DETAILED DESCRIPTION

[0058] The following describes the embodiments of the present invention through specific embodiments, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0059] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention. The illustrations only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0060] The corpus management method and system of the present invention can select the corresponding collection channel according to the corpus requirements to obtain the original corpus, and after obtaining the original corpus, pre-process the original corpus to improve the subsequent processing efficiency; in the subsequent processing process, the original corpus is screened again by corpus cleaning and screening to obtain the target corpus to improve the corpus quality, and the target corpus is annotated to obtain annotated data, which is convenient for subsequent users to understand when using the target corpus, and improves the efficiency and convenience of users using the target corpus, and according to different formats of the target corpus, different target corpora and corresponding annotated data are packaged to obtain a packaged data packet, which is convenient for users to call and ensures the security of the corpus; and the packaged data packet is maintained during the entire use cycle, which can effectively extend the service life of the corpus.

[0061] The storage medium of the present invention stores a computer program, and when the computer program is executed by a processor, the following corpus management method is implemented. The storage medium includes: a read-only memory (ROM), a random access memory (RAM), a disk, a USB flash drive, a memory card, or an optical disk, etc., which can store program codes.

[0062] Any combination of one or more storage media may be used. The storage medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, device, or device.

[0063] Computer-readable signal media may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0064] The program code embodied on the computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0065] Computer program code for performing the operation of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0066] The present invention will be described below with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present invention. It should be understood that each box of the flowchart and / or block diagram and the combination of the boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer or other programmable data processing device, so as to produce a machine, so that when these computer program instructions are executed by the processor of the computer or other programmable data processing device, a device for implementing the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated.

[0067] These computer program instructions may also be stored in a computer-readable medium, which enables a computer, other programmable data processing apparatus, or other device to operate in a specific manner, so that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0068] Computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0069] The terminal of the present invention includes a processor and a memory.

[0070] The memory is used to store computer programs; preferably, the memory includes: ROM, RAM, disk, USB flash drive, memory card or CD, etc., which can store program codes.

[0071] The processor is connected to the memory and is used to execute the computer program stored in the memory so that the terminal executes the following corpus management method.

[0072] Preferably, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0073] like Figure 1 As shown, in one embodiment, the present invention discloses a corpus management method, comprising the following steps:

[0074] S100: According to the corpus requirement, select a corresponding collection channel to obtain original corpus, and pre-process the original corpus to obtain intermediate corpus.

[0075] In some embodiments, the collection channel includes a public data source channel, a private procurement channel, and a shared collection channel. According to the corpus requirements, selecting a corresponding collection channel to obtain the original corpus includes:

[0076] Obtaining the user's corpus requirement, and determining the type of the required original corpus according to the corpus requirement;

[0077] Selecting a corresponding collection channel according to the source type of the original corpus to obtain the original corpus;

[0078] Wherein, when the source type of the original corpus is a public type, selecting a public data source channel to obtain the original corpus;

[0079] When the source type of the original corpus is a private type, selecting the private purchasing channel to obtain the original corpus;

[0080] When the source type of the original corpus is a shared type, the shared collection channel is selected to obtain the original corpus.

[0081] In this embodiment, after obtaining the user's corpus requirements, different collection channels are selected to obtain the original corpus according to different corpus requirements, thereby ensuring the rationality and accuracy of the original corpus source and avoiding the problem of non-compliance of the original corpus.

[0082] Specifically, the corpus requirement includes the source type of the corpus, and the source type includes a public type, a private type, and a shared type. When the source type of the original corpus is a public type, a public data source channel is selected to obtain the original corpus;

[0083] When the source type of the original corpus is a private type, the private procurement channel is selected to obtain the original corpus; when the source type of the original corpus is a shared type, the shared acquisition channel is selected to obtain the original corpus.

[0084] Furthermore, for the acquisition of public corpus, first determine the legally public data sources, including government public data, academic paper data, public social media data, public data sets, etc. Classify the data sources, such as text data, image data, video data, etc., and record their main characteristics and scope of application. And conduct compliance checks on the data, check the use license and copyright statement of the data source, and ensure that the acquisition and use of the data comply with relevant laws and regulations and platform terms of use. For data sources with open licenses, such as data licensed by Creative Commons, confirm their specific licensing terms and ensure that relevant regulations are followed during use. In the process of data collection, data collection technical specifications, formulate technical specifications for data collection, including the use of data crawling tools, data format conversion, data storage, etc., to ensure information security during data collection and avoid illegal intrusion and data leakage.

[0085] At the same time, a data list should be established to record in detail the acquisition time, data volume, data type, permission status and other information of each public data source, and the data list should be updated regularly to ensure the timeliness and accuracy of the data source.

[0086] For the acquisition of private corpus, the first step is to select suppliers. After determining that the data quality, service capabilities, and credibility meet the requirements, a suitable data supplier is selected. At the same time, data acceptance standards are formulated, including data completeness, accuracy, and timeliness, and suppliers are evaluated regularly to ensure the continued stability and high quality of their data services.

[0087] On the other hand, for the acquisition of shared corpus, we select sharing partners that meet the requirements by evaluating their resource instructions, technical capabilities and data management levels.

[0088] The cooperation model with the shared partners includes data sharing, joint research, and data exchange, so as to flexibly respond to the needs of different partners.

[0089] In some embodiments, preprocessing the original corpus to obtain the intermediate corpus includes:

[0090] Copy the original corpus to obtain a first control corpus and a second control corpus, evenly split the original corpus into N basic corpus groups in order, randomly evenly split the first control corpus into N first corpus groups, and randomly evenly split the second control corpus group into N second corpus groups, where N is a positive integer;

[0091] Respectively obtaining data completeness, data accuracy, and data timeliness of N basic corpus groups, N first corpus groups, and N second corpus groups, and respectively obtaining quality evaluation values ​​of the N basic corpus groups, first evaluation values ​​of the N first corpus groups, and second evaluation values ​​of the N second corpus groups according to the corresponding data completeness, data accuracy, and data timeliness;

[0092] Calculating a first mean of the first evaluation values ​​of the N first corpus groups and a second mean of the second evaluation values ​​of the N second corpus groups, and calculating a standard average value based on the first mean and the second mean;

[0093] The quality evaluation value is compared with the standard average value, and the basic corpus groups corresponding to the quality evaluation values ​​smaller than the standard average value are eliminated, and the remaining basic corpus groups are recombined to obtain the intermediate corpus.

[0094] In this embodiment, in order to ensure the quality of the corpus obtained subsequently, the original corpus is preprocessed to obtain the intermediate corpus, thereby reducing the corpus that does not meet the requirements.

[0095] Specifically, after obtaining the original corpus, the original corpus is evenly divided into multiple basic corpus groups of the same size in order, and the original corpus is copied to obtain the first control corpus and the second control corpus, and randomly divided into N first corpus groups and N second corpus groups, respectively, and then the data completeness, data accuracy and data timeliness of the N basic corpus groups, N first corpus groups and N second corpus groups are obtained respectively. The data completeness is obtained by calculating the proportion of missing values ​​in each corpus group, the data accuracy is obtained by calculating the numerical distribution in each corpus group, and the data timeliness is obtained by calculating the data update time in each corpus group. Since the calculation process of the above data completeness, the data accuracy and the data timeliness are all the contents of the prior art, the present invention does not involve its own improvement, and it will not be repeated here.

[0096] On this basis, the quality evaluation value of the basic corpus group, the first evaluation value of the first corpus group, and the second evaluation value of the second corpus group are calculated according to the data completeness, data accuracy, and data timeliness.

[0097] The calculation process of the quality assessment value S satisfies the following formula:

[0098] , where A, B, and C represent data completeness, data accuracy, and data timeliness, respectively; a, b, and c represent weight values ​​of data completeness, data accuracy, and data timeliness, respectively; the weight values ​​can be manually set empirical values ​​or obtained through weight calculation methods such as the hierarchical analysis method, which will not be described here.

[0099] Similarly, the calculation process of the first evaluation value and the second evaluation value is the same as the calculation process of the quality evaluation value, which will not be repeated here.

[0100] S200: According to a preset standard, the intermediate corpus is cleaned to obtain a cleaned corpus.

[0101] In some other embodiments, the step of performing corpus cleaning on the intermediate corpus according to a preset standard to obtain a cleaned corpus includes:

[0102] Determine the target format according to the preset standard, and convert the intermediate corpus into the target format;

[0103] Performing data cleaning on the intermediate corpus after the format conversion, and performing normalization adjustment after the data cleaning to obtain the first corpus;

[0104] After completing the normalization adjustment, performing balancing processing on the first corpus to obtain a second corpus;

[0105] Segment and deduplicate the second corpus to obtain a cleaned corpus, where each of the cleaned corpora includes multiple data sets.

[0106] In this embodiment, during the cleaning process of the intermediate corpus, first determine the corresponding target format according to a preset standard, such as a text file, a CSV file, or a database format, convert the intermediate corpus into the corresponding target format, and then perform data cleaning on the intermediate corpus, including noise removal, stop word processing, and abnormal data processing.

[0107] Among them, noise removal is mainly used to remove noise information such as invalid characters, garbled codes, and repeated punctuation marks in the corpus. Stop word processing mainly removes common meaningless words in the corpus through a stop word list, such as words like "de" and the final "le". Abnormal data processing is to detect and process abnormal data in the corpus, such as overly long or short texts and missing values.

[0108] After data cleaning, perform normalization adjustment on the intermediate corpus to obtain the first corpus. The normalization adjustment includes text standardization, punctuation normalization, and vocabulary normalization. Text standardization includes unifying the case, full-width and half-width characters, and traditional and simplified character conversions in the text. Punctuation normalization includes uniformly using the standard punctuation format and removing redundant spaces, line breaks, etc. Vocabulary normalization includes uniformly replacing synonyms and near-synonyms to ensure the consistency of the corpus.

[0109] After obtaining the first corpus, perform balancing processing on the first corpus to obtain the second corpus. The balancing processing mainly targets class-imbalanced data and performs oversampling or undersampling to ensure the balance of data in each class.

[0110] After obtaining the second corpus, segment and deduplicate the second corpus to obtain a cleaned corpus.

[0111] Specifically, the segmentation and removal include data deduplication and data segmentation. When performing data deduplication, remove duplicate corpora through a hash algorithm or a text similarity calculation method. Data segmentation is to divide the corpus into multiple subsets according to the nature and use of the corpus, including a training set, a validation set, and a test set.

[0112] During the process of corpus cleaning, corresponding cleaning standards can be formulated. For example, there are no obvious missing, broken sentences, or garbled codes in the corpus to ensure the integrity of the corpus; the punctuation and vocabulary usage in the corpus should be consistent to avoid noise caused by excessive diversity; the information in the corpus should be true and accurate to avoid false information and misleading content; the corpus should follow language norms and use standard grammar, vocabulary, and formats to improve the readability and usability of the corpus.

[0113] S300: According to the value alignment rule, the cleaned corpus is aligned and screened to obtain the target corpus.

[0114] In some other embodiments, the aligning and screening the cleaned corpus according to the value alignment rule to obtain the target corpus includes:

[0115] The cleaned corpus is subjected to ethical screening, legal screening and commercial screening in turn, and the corpus that does not meet the ethical requirements, legal requirements and commercial requirements is removed according to the screening results to obtain the target corpus.

[0116] Among them, the ethical screening is to ensure that the corpus and model outputs comply with social ethical norms and avoid discrimination, bias and misleading information. Promote the fairness and impartiality of models and corpora and avoid unfair algorithmic decisions. Legal screening is to ensure that the collection, processing and use of corpora comply with relevant laws and regulations, such as data protection laws and privacy laws, to avoid corpora and model outputs infringing on the privacy and intellectual property rights of others. Commercial screening is to ensure that the application of corpora can bring commercial value to enterprises and support the strategic goals and market needs of enterprises. Balance short-term interests and long-term value to ensure the sustainable use of corpora.

[0117] Exemplarily, when performing alignment screening, the specific process includes:

[0118] (1) Demand analysis:

[0119] Communicate with business departments and relevant stakeholders to clarify the application scenarios and value requirements of the corpus.

[0120] Analyze business needs and determine key value indicators such as customer satisfaction, market share, brand image, etc.

[0121] (2) Value Assessment:

[0122] Conduct value assessment on existing corpus and evaluate its potential value and risks in business applications.

[0123] Use a multi-dimensional evaluation method, including quantitative analysis and qualitative analysis, to comprehensively evaluate the value of the corpus.

[0124] (3) Value screening:

[0125] Based on the value assessment results, the corpus is screened, high-value corpus is retained, and low-value or high-risk corpus is eliminated.

[0126] Establish value screening criteria and processes to ensure the fairness and transparency of the screening process.

[0127] (4) Value optimization:

[0128] Optimize the selected corpus, such as data enhancement, denoising, synthesis, etc., to enhance the application value of the corpus. Introduce advanced technologies, such as natural language processing and machine learning, to improve the standards of intelligent and automated processing of corpus.

[0129] The standards for ethical screening and commercial screening are formulated according to usage requirements, while the standards for legal screening are formulated according to existing laws. The above-mentioned process of corpus alignment screening is the content of the prior art and will not be repeated here.

[0130] S400: Annotate the target corpus to obtain annotation data corresponding to each target corpus.

[0131] In some embodiments, the tagging of the target corpus to obtain the tagging data corresponding to each target corpus includes:

[0132] Classifying the target corpus according to the data type of the target corpus to obtain classified corpus, and formulating corresponding annotation specifications for the classified corpus;

[0133] Performing data preprocessing on each of the classified corpora according to the annotation specification and then segmenting the data to obtain multiple data sets;

[0134] Preliminarily annotating each of the data sets to obtain primary annotated data;

[0135] The primary annotated data are reviewed and feedback is provided, the primary annotated data are modified and adjusted according to the feedback, and the adjusted primary annotated data are combined together to obtain the annotated data corresponding to the target corpus.

[0136] Corpus annotation is a crucial part of the large model training process. Its purpose is to add semantic information to the original data so that the model can learn and understand the deep meaning of the data. Corpus annotation is an important part of the preparation of large model training data. A complete rule system is crucial to improving the quality and efficiency of annotation. By standardizing the annotation process, formulating annotation standards, and adopting advanced technologies and management measures, the accuracy and consistency of corpus annotation can be effectively improved, providing high-quality annotated data for large model training.

[0137] In this embodiment, when the target corpus is annotated, the target corpus is first classified according to the data type of the target corpus, including text type, image type, audio type and video type, to determine the pertinence and effectiveness of the annotation process, and corresponding annotation specifications are formulated for different types of data to facilitate the subsequent annotation of different types of corpus. Before annotation, the annotation task is defined to clarify the goals and requirements of each annotation task, such as text classification, entity recognition, etc.

[0138] Among them, the annotation specifications include annotation standards, consistency requirements, accuracy requirements and timeliness requirements. The annotation standards include annotation rules, such as category definitions for named entity recognition and classification standards for sentiment analysis, which are easy for annotators to understand and execute. The consistency requirements are mainly used to ensure that the annotation results of the same data at different locations and times are consistent, avoiding the situation where the same data has different annotations. The accuracy requirements are to ensure accurate annotation, and the timeliness requirements are to clarify the annotation time of the data to avoid affecting the annotation progress.

[0139] After formulating the annotation specifications, different classified corpora are preprocessed and segmented to obtain multiple data sets. The preprocessing process includes format conversion, data cleaning, sample sampling, etc. to ensure that the data is suitable for the annotation task. By segmenting the classified corpora to obtain multiple data sets, data leakage and annotation bias can be effectively avoided, the diversity and representativeness of the annotation data can be ensured, and the accuracy of corpus annotation can be improved.

[0140] After that, each data set is preliminarily labeled to obtain preliminary labeled data, and then the preliminary labeled data is self-checked and cross-checked to ensure the consistency and accuracy of the labeling results, and further reviewed by sampling inspection to ensure the accuracy of the labeling.

[0141] In the process of reviewing the primary annotated data, the review opinions are promptly fed back so that the primary annotated data can be subsequently modified and adjusted according to the review opinions to obtain the adjusted and optimized annotated data, and the annotated data of different data sets can be combined to obtain the annotated data corresponding to the target corpus.

[0142] Among them, a corresponding database is established for the feedback of primary annotation data to facilitate subsequent query of the feedback results.

[0143] After reviewing and providing feedback on the primary annotated data, the primary annotated data are modified and adjusted according to the feedback, and the adjusted primary annotated data are combined together to obtain the annotated data corresponding to the target corpus.

[0144] S500: According to the format of the target corpus, encapsulate the target corpus and the annotated data together to obtain an encapsulated data packet, and maintain the encapsulated data packet.

[0145] Corpus encapsulation is the process of organizing, classifying and formatting the cleaned and annotated corpus to facilitate storage, transmission and use. Corpus encapsulation is an important link to ensure that data can be efficiently used and managed. A complete encapsulation rule system is crucial to improving the quality and use value of data. In this embodiment, by standardizing the encapsulation process, formulating encapsulation standards and formats, and strengthening data security and management, the efficiency, reliability and security of corpus encapsulation can be ensured, providing a solid data foundation for the training and application of large models.

[0146] In some embodiments, encapsulating the target corpus and the annotated data together to obtain an encapsulated data packet according to the format of the target corpus includes:

[0147] After arranging and backing up the classified corpora, determining a packaging standard and packaging format for each of the classified corpora;

[0148] Encapsulating the classified corpus and the annotated data according to the encapsulation standard and the encapsulation format to obtain encapsulated data;

[0149] According to the set metadata structure, metadata corresponding to the packaged data is generated during the packaging process, and the metadata is used to record relevant information of the packaged data, including data source, collection time, annotation information and usage authority;

[0150] The encapsulated data and the metadata are combined together and compressed and encrypted to form an encapsulated data packet.

[0151] Specifically, the classified corpus is first sorted and backed up to ensure data integrity and consistency, and to prevent data loss or damage that cannot be retrieved during the packaging process through backup. Then the packaging standard and packaging format of each classified corpus are determined to facilitate subsequent packaging.

[0152] The packaging standards for text data include file format, structured format, and encoding format. The file format uniformly uses standardized text file formats, such as TXT, CSV, JSON, etc. The structured format ensures the structuring of text data, such as each row represents a sample, and each column represents a feature or label. The encoding format uniformly uses UTF-8 encoding to avoid reading errors caused by inconsistent encoding.

[0153] Image data packaging includes file format, directory structure, and file naming. Common image file formats are used, such as JPEG, PNG, TIFF, etc. Directory structure: Establish a directory structure according to category or purpose to ensure the orderly storage of image files. File naming: Use a unified naming rule, including category, number and other information, to facilitate retrieval and management.

[0154] Audio data packaging includes file format, directory structure, and file naming. The file format uses standard audio file formats, such as WAV, MP3, FLAC, etc. The directory structure is established according to category, language or purpose to ensure the orderly storage of audio files. File naming uses a unified naming rule, including category, number, recording date and other information.

[0155] Video data packaging includes file format, directory structure, and file naming. File format: Use standard video file formats, such as MP4, AVI, MKV, etc. Directory structure: Establish directory structure according to category, content or purpose to ensure orderly storage of video files. File naming: Use a unified naming rule, including category, number, recording date and other information.

[0156] After determining the encapsulation standard and encapsulation format, the target corpus and the annotated data are encapsulated together to obtain encapsulated data. Then, according to the set metadata structure, metadata corresponding to the encapsulated data is generated during the encapsulation process, and the metadata is used to record relevant information of the encapsulated data, including data source, collection time, annotation information and usage rights.

[0157] Specifically, metadata is used to record relevant information of the encapsulated data, including data source, collection time, annotation information, usage rights, etc., and the metadata adopts a standard format and structure for easy storage and retrieval. During the encapsulation process, the metadata information of each encapsulated data is fully recorded to ensure the traceability and manageability of the data package. At the same time, the metadata file is stored together with the encapsulated data to ensure the consistency of the data and metadata.

[0158] Afterwards, the encapsulated data and the metadata are combined together and compressed and encrypted to form an encapsulated data packet, completing the encapsulation of the encapsulated data and metadata, which is convenient for subsequent tracing and management.

[0159] In some embodiments, in the process of generating the encapsulated data packet, the method further includes:

[0160] Writing a packaging description document, recording relevant parameters of the packaging data packet through the packaging description document;

[0161] Generate an encapsulation log to record the operation content of the encapsulated data packet through the encapsulation log;

[0162] Among them, the relevant parameters include the directory structure, file naming rules, compression and encryption methods, encapsulation standards and formats, metadata definitions and encapsulation processes of the encapsulated data, and the encapsulation log includes encapsulation steps, operation logs, operation time, operators and operation content.

[0163] Furthermore, in the process of generating the encapsulated data package, an encapsulation description document is further compiled to describe in detail the encapsulation process, encapsulation standards and formats, metadata definitions and records, etc. The description document includes the directory structure, file naming rules, compression and encryption methods, etc. of the encapsulated data, which facilitates users to understand and use the encapsulated data. And generate an encapsulation log to record the key steps and operation logs in the encapsulation process, such as data collation, compression, encryption, etc. The encapsulation log should include information such as operation time, operator, and operation content to ensure the transparency and traceability of the encapsulation process.

[0164] In some other embodiments, maintaining the encapsulated data packet includes:

[0165] Acquire usage information and feedback information of the target corpus in each of the encapsulated data packets, and calculate the value parameter of the target corpus according to the usage information and the feedback information;

[0166] Regularly adjusting the order of the target corpus in the encapsulated data packet according to the value parameter, and regenerating a new encapsulated data packet according to the adjusted target corpus;

[0167] Adjusting the storage period of the target corpus according to the value parameter;

[0168] The greater the value parameter of the target corpus is, the longer the corresponding adjusted storage period is.

[0169] In this embodiment, during the process of the user using the encapsulated data package, the usage information and feedback information of the target corpus in the encapsulated data package is collected so as to calculate the value parameter of the target corpus based on the usage information and feedback information. Then, the target corpus in the encapsulated data package is adjusted based on the value parameter, and the adjusted target corpus is re-annotated, and a new encapsulated data package is generated. In addition, the storage period of the target corpus is adjusted based on the value parameter. The target corpus with a low value parameter has a shorter storage period, and the target corpus with a high value parameter has a longer storage period.

[0170] Furthermore, the usage information includes a usage rate, the feedback information includes a positive feedback rate and a negative feedback rate, the value parameter is positively correlated with the usage rate and the positive feedback rate, and is negatively correlated with the negative feedback rate. The specific calculation process adopts the content of the existing technology and will not be repeated here.

[0171] When adjusting the target corpus in the encapsulated data package according to the value parameter, the order of the target corpus with a large value parameter is adjusted forward, and the order of the target corpus with a small value parameter is adjusted backward, so as to facilitate the user to quickly retrieve the high-value corpus when using it, thereby ensuring the efficiency of the use of the target corpus after subsequent adjustment and enhancing the user experience.

[0172] Among them, the adjustment period of the target corpus in the encapsulated data package is manually set according to the value parameter, which can be selected according to actual needs and will not be repeated here.

[0173] It should be noted that the protection scope of the corpus management method described in the present invention is not limited to the execution order of the steps listed in this embodiment. All solutions implemented by adding, reducing or replacing steps in the prior art based on the principles of the present invention are included in the protection scope of the present invention.

[0174] like Figure 2 As shown, in one embodiment, the present invention provides a corpus management system, referring to Figure 2 ,include:

[0175] The preprocessing module 201 is used to select a corresponding collection channel to obtain original corpus according to the corpus requirements, and preprocess the original corpus to obtain intermediate corpus;

[0176] A cleaning module 202, configured to clean the intermediate corpus according to a preset standard to obtain a cleaned corpus;

[0177] A screening module 203 is used to perform alignment screening on the cleaned corpus according to the value alignment rule to obtain the target corpus;

[0178] The annotation module 204 is used to annotate the target corpus to obtain annotation data corresponding to each target corpus;

[0179] The encapsulation maintenance module 205 is used to encapsulate the target corpus and the annotated data together according to the format of the target corpus to obtain an encapsulated data packet, and maintain the encapsulated data packet.

[0180] It should be noted that the structure and principle of the above-mentioned corpus management system correspond one-to-one to the steps in the above-mentioned corpus management method, so they will not be repeated here.

[0181] It should be noted that it should be understood that the division of the various modules of the above system is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also be all implemented in the form of hardware; some modules can also be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the x module can be a separately established processing element, or it can be integrated in a certain chip of the above system for implementation. In addition, it can also be stored in the memory of the above system in the form of program code, and called and executed by a certain processing element of the above system. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each module above can be completed by an integrated logic circuit of hardware in the processor element or instructions in the form of software.

[0182] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASIC), or one or more digital signal processors (DSP), or one or more field programmable gate arrays (FPGA). For another example, when a module above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0183] It should be noted that the corpus management system of the present invention can implement the corpus management method of the present invention, but the implementation device of the corpus management method of the present invention includes but is not limited to the structure of the corpus management system listed in this embodiment, and all structural deformations and replacements of the prior art made according to the principles of the present invention are included in the protection scope of the present invention.

[0184] In summary, the corpus management method and system of the present invention can select the corresponding collection channel according to the corpus requirements to obtain the original corpus, and after obtaining the original corpus, pre-process the original corpus to improve the efficiency of subsequent processing; in the subsequent processing process, the original corpus is screened again by corpus cleaning and screening to obtain the target corpus to improve the quality of the corpus, and the target corpus is annotated to obtain annotated data, which is convenient for subsequent users to understand when using the target corpus, thereby improving the efficiency and convenience of users using the target corpus, and according to the different formats of the target corpus, different target corpora and corresponding annotated data are packaged to obtain packaged data packets, which is convenient for users to call and ensures the security of the corpus; and the packaged data packets are maintained throughout the entire use cycle, which can effectively extend the service life of the corpus; therefore, the present invention effectively overcomes the various shortcomings in the prior art and has a high industrial utilization value.

[0185] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.

Claims

1. A corpus management method, characterized in that: include: According to the corpus requirements, select the corresponding collection channel to obtain the original corpus, and preprocess the original corpus to obtain the intermediate corpus; According to a preset standard, the intermediate corpus is cleaned to obtain a cleaned corpus; According to the value alignment rule, the cleaned corpus is aligned and screened to obtain the target corpus; Performing corpus annotation on the target corpus to obtain annotation data corresponding to each target corpus; According to the format of the target corpus, encapsulate the target corpus and the annotated data together to obtain an encapsulated data packet, and maintain the encapsulated data packet; The preprocessing of the original corpus to obtain the intermediate corpus includes: Copy the original corpus to obtain a first control corpus and a second control corpus, evenly split the original corpus into N basic corpus groups in order, randomly evenly split the first control corpus into N first corpus groups, and randomly evenly split the second control corpus group into N second corpus groups, where N is a positive integer; Respectively obtaining data completeness, data accuracy, and data timeliness of N basic corpus groups, N first corpus groups, and N second corpus groups, and respectively obtaining quality evaluation values ​​of the N basic corpus groups, first evaluation values ​​of the N first corpus groups, and second evaluation values ​​of the N second corpus groups according to the corresponding data completeness, data accuracy, and data timeliness; Calculating a first mean of the first evaluation values ​​of the N first corpus groups and a second mean of the second evaluation values ​​of the N second corpus groups, and calculating a standard average value based on the first mean and the second mean; The quality evaluation value is compared with the standard average value, and the basic corpus groups corresponding to the quality evaluation values ​​smaller than the standard average value are eliminated, and the remaining basic corpus groups are recombined to obtain the intermediate corpus.

2. The corpus management method according to claim 1, characterized in that: The collection channels include public data source channels, private procurement channels and shared collection channels. According to the corpus requirements, the corresponding collection channels are selected to obtain the original corpus, including: Obtaining the user's corpus requirement, and determining the type of the required original corpus according to the corpus requirement; Selecting a corresponding collection channel according to the source type of the original corpus to obtain the original corpus; Wherein, when the source type of the original corpus is a public type, selecting a public data source channel to obtain the original corpus; When the source type of the original corpus is a private type, the private procurement channel is selected to obtain the original corpus; when the source type of the original corpus is a shared type, the shared acquisition channel is selected to obtain the original corpus.

3. The corpus management method according to claim 1, characterized in that: The step of performing corpus cleaning on the intermediate corpus according to a preset standard to obtain a cleaned corpus includes: Determine the target format according to the preset standard, and convert the intermediate corpus into the target format; Performing data cleaning on the intermediate corpus after the format conversion, and performing normalization adjustment after the data cleaning to obtain the first corpus; After completing the normalization adjustment, performing balancing processing on the first corpus to obtain a second corpus; The second corpus is segmented and deduplicated to obtain cleaned corpuses, wherein each of the cleaned corpuses includes multiple data sets.

4. The corpus management method according to claim 1, characterized in that: The step of aligning and screening the cleaned corpus according to the value alignment rule to obtain the target corpus includes: The cleaned corpus is subjected to ethical screening, legal screening and commercial screening in turn, and the corpus that does not meet the ethical requirements, legal requirements and commercial requirements is removed according to the screening results to obtain the target corpus.

5. The corpus management method according to claim 1, characterized in that: The tagging of the target corpus to obtain the tagging data corresponding to each target corpus includes: Classifying the target corpus according to the data type of the target corpus to obtain classified corpus, and formulating corresponding annotation specifications for the classified corpus; Performing data preprocessing on each of the classified corpora according to the annotation specification and then segmenting the data to obtain multiple data sets; Preliminarily annotating each of the data sets to obtain primary annotated data; The primary annotated data are reviewed and feedback is provided, the primary annotated data are modified and adjusted according to the feedback, and the adjusted primary annotated data are combined together to obtain the annotated data corresponding to the target corpus.

6. The corpus management method according to claim 5, characterized in that: The step of encapsulating the target corpus and the annotated data together according to the format of the target corpus to obtain an encapsulated data packet includes: After arranging and backing up the classified corpora, determining a packaging standard and packaging format for each of the classified corpora; Encapsulating the classified corpus and the annotated data according to the encapsulation standard and the encapsulation format to obtain encapsulated data; According to the set metadata structure, metadata corresponding to the packaged data is generated during the packaging process, and the metadata is used to record relevant information of the packaged data, including data source, collection time, annotation information and usage authority; The encapsulated data and the metadata are combined together and compressed and encrypted to form an encapsulated data packet.

7. The corpus management method according to claim 6, characterized in that: In the process of generating the encapsulated data packet, the method further comprises: Writing a packaging description document, recording relevant parameters of the packaging data packet through the packaging description document; Generate an encapsulation log to record the operation content of the encapsulated data packet through the encapsulation log; Among them, the relevant parameters include the directory structure, file naming rules, compression and encryption methods, encapsulation standards and formats, metadata definitions and encapsulation processes of the encapsulated data, and the encapsulation log includes encapsulation steps, operation logs, operation time, operators and operation content.

8. The corpus management method according to claim 6, characterized in that: The maintaining of the encapsulated data packet comprises: Acquire usage information and feedback information of the target corpus in each of the encapsulated data packets, and calculate the value parameter of the target corpus according to the usage information and the feedback information; Regularly adjusting the order of the target corpus in the encapsulated data packet according to the value parameter, and regenerating a new encapsulated data packet according to the adjusted target corpus; Adjusting the storage period of the target corpus according to the value parameter; The greater the value parameter of the target corpus is, the longer the corresponding adjusted storage period is.

9. A corpus management system, characterized in that: include: A preprocessing module is used to select a corresponding collection channel to obtain original corpus according to corpus requirements, and preprocess the original corpus to obtain intermediate corpus; A cleaning module, used for performing corpus cleaning on the intermediate corpus according to a preset standard to obtain a cleaned corpus; A screening module, used for performing alignment screening on the cleaned corpus according to a value alignment rule to obtain a target corpus; A tagging module, used for tagging the target corpus to obtain tagging data corresponding to each target corpus; A packaging and maintenance module, used for packaging the target corpus and the annotated data together according to the format of the target corpus to obtain a packaged data packet, and maintaining the packaged data packet; The preprocessing of the original corpus to obtain the intermediate corpus includes: Copy the original corpus to obtain a first control corpus and a second control corpus, evenly split the original corpus into N basic corpus groups in order, randomly evenly split the first control corpus into N first corpus groups, and randomly evenly split the second control corpus group into N second corpus groups, where N is a positive integer; Respectively obtaining data completeness, data accuracy, and data timeliness of N basic corpus groups, N first corpus groups, and N second corpus groups, and respectively obtaining quality evaluation values ​​of the N basic corpus groups, first evaluation values ​​of the N first corpus groups, and second evaluation values ​​of the N second corpus groups according to the corresponding data completeness, data accuracy, and data timeliness; Calculating a first mean of the first evaluation values ​​of the N first corpus groups and a second mean of the second evaluation values ​​of the N second corpus groups, and calculating a standard average value based on the first mean and the second mean; The quality evaluation value is compared with the standard average value, and the basic corpus groups corresponding to the quality evaluation values ​​smaller than the standard average value are eliminated, and the remaining basic corpus groups are recombined to obtain the intermediate corpus.