Corpus Directory Management Method and System for Industrial Large Models
By preprocessing and scene analysis of the target corpus, calculating standard directory values, and generating target catalogs, the problem of low corpus storage and call efficiency is solved, and rapid query and storage is achieved.
Patent Information
- Application Number
- CN202510380061.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-28
AI Technical Summary
In the prior art, corpus storage and call efficiency are low, resulting in inconvenient corpus management.
By obtaining the basic information of the target corpus for preprocessing, performing scenario analysis to obtain label information, and calculating standard directory values, generating target directory for storage.
It improves the storage and calling efficiency of corpus, and realizes rapid query and storage of corpus.
Smart Images

Figure CN119903124B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of big data, and particularly relates to a method and system for managing a corpus directory for an industrial large model. Background Art
[0002] An artificial intelligence large model refers to a "large parameter" model trained using large-scale data and powerful computing capabilities. These models usually have high generality and generalization capabilities and can be applied to fields such as natural language processing, image recognition, and speech recognition. They can be classified into large language models, visual large models, multi-modal large models, and basic large models.
[0003] And a corpus is an important material for training an artificial intelligence large model. Generally, a corpus refers to instances and data sets used in linguistic research and natural language processing. It can be written text, spoken records, or other structured data, and is usually used for analyzing language phenomena, supporting machine translation, speech recognition, automatic text summarization, and other tasks. Usually, a corpus is a set of text or speech data that has been collected, sorted, and annotated. These data are used in linguistic research to analyze language usage rules, vocabulary changes, and grammatical structures, etc. In natural language processing, a corpus is the basic data source for training and testing models, supporting functions such as machine translation, speech recognition, and sentiment analysis.
[0004] Due to the extremely large number of corpora, they are generally stored in a classified storage manner, and when storing, a multi-level directory is established for storage. When storing and searching, the corresponding directory needs to be determined layer by layer. Due to the refined classification of current corpora, the directory content of many corpora is too much, resulting in low efficiency in storing and searching for corpora, which is not conducive to the storage and invocation of corpora. Summary of the Invention
[0005] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method and system for managing a corpus directory for an industrial large model, which is used to solve the problem of low efficiency in storing and invoking corpora in the prior art.
[0006] To achieve the above purpose and other related purposes, the present invention provides a method for managing a corpus directory for an industrial large model, including the following steps:
[0007] Obtain the basic information of the target corpus, and preprocess the target corpus according to the basic information to obtain a standard corpus;
[0008] Perform scenario analysis on the standard corpus to obtain corresponding tag information;
[0009] Calculate the directory value of the standard corpus according to the tag information to obtain the standard directory value of the standard corpus;
[0010] A corresponding target directory is generated according to the standard directory value, and the standard corpus is mapped to the target directory and then stored.
[0011] In some embodiments, preprocessing the target corpus according to the basic information to obtain a standard corpus includes:
[0012] Performing field scanning on the target corpus to remove invalid fields and noise fields in the target corpus to obtain a first corpus;
[0013] Obtaining a sensitive word library corresponding to the target corpus based on the basic information, comparing the first corpus with the sensitive word library, and replacing sensitive words in the first corpus with label words to obtain a corresponding standard corpus;
[0014] The label words correspond one-to-one to each of the sensitive words.
[0015] In some embodiments, the basic information includes at least industry information, modality information, and language information, and the performing scenario analysis on the standard corpus to obtain corresponding tag information includes:
[0016] Acquire the industry type of the target corpus according to the industry information, convert the industry type into a preset coding format, and generate a first label matrix corresponding to the standard corpus;
[0017] Acquiring a modality type in the target corpus according to the modality information, converting the modality type into a preset encoding format, and generating a second label matrix corresponding to the standard corpus;
[0018] Acquiring a language type of the target corpus according to the language information, converting the language type into a preset encoding format, and generating a third label matrix corresponding to the standard corpus;
[0019] Integrating the first label matrix, the second label matrix, and the third label matrix in sequence into a label matrix to obtain the label information;
[0020] The rows and columns of the first label matrix, the second label matrix and the third label matrix are all the same.
[0021] In some embodiments, calculating the catalog value of the standard corpus according to the tag information to obtain the standard catalog value of the standard corpus includes:
[0022] Calculating the same-position average of the first label matrix, the second label matrix, and the third label matrix to obtain an average label matrix;
[0023] Obtain the corpus sequence of the standard corpus and generate the corpus index value of the corpus sequence;
[0024] Calculate the ratios between the eigenvalues corresponding to the first label matrix, the second label matrix, the third label matrix, and the average label matrix and the corpus index value respectively to obtain a first catalog value, a second catalog value, a third catalog value, and an average catalog value respectively;
[0025] Combine the first catalog value, the second catalog value, the third catalog value, and the average catalog value together to generate a catalog sequence;
[0026] Calculate the catalog index value of the catalog sequence, determine the standard catalog value of the standard corpus according to the corpus index value and the catalog index value, and establish a mapping relationship between the standard catalog value and the standard corpus.
[0027] In some embodiments, the generating a corresponding target catalog according to the standard catalog value and storing the standard corpus after establishing a mapping between the standard corpus and the target catalog includes:
[0028] Correspondingly obtain the corresponding catalog path in the catalog library according to the standard catalog value, and generate a target catalog corresponding to the standard corpus according to the catalog path;
[0029] After establishing a mapping between the standard corpus and the target catalog, store the standard corpus in the target catalog.
[0030] In some embodiments, the correspondingly obtaining the corresponding catalog path in the catalog library according to the standard catalog value and generating a target catalog corresponding to the standard corpus according to the catalog path includes:
[0031] Split the standard catalog value into a plurality of sub-catalog values corresponding to the number, and generate sub-paths in order according to the plurality of sub-catalog values;
[0032] Combine the plurality of sub-paths together to obtain the catalog path, and generate a corresponding target catalog according to the catalog path.
[0033] In some embodiments, the combining the first catalog value, the second catalog value, the third catalog value, and the average catalog value together to generate a catalog sequence includes:
[0034] Insert delimiters at the head and tail ends of the first catalog value, the second catalog value, the third catalog value, and the average catalog value respectively, and then arrange them in order to generate the catalog sequence.
[0035] In some embodiments, the standard catalog value of the standard corpus is a sequence combination value of the corpus index value and the catalog index value.
[0036] The present invention also provides a corpus directory management system for industrial large models, including:
[0037] A preprocessing module, configured to obtain the basic information of the target corpus and preprocess the target corpus according to the basic information to obtain a standard corpus;
[0038] An analysis module, configured to perform scenario analysis on the standard corpus to obtain corresponding tag information;
[0039] A calculation module, configured to calculate the directory value of the standard corpus according to the tag information to obtain the standard directory value of the standard corpus;
[0040] A directory generation module, configured to generate a corresponding target directory according to the standard directory value, establish a mapping between the standard corpus and the target directory, and then store them.
[0041] As described above, the corpus directory management method and system for industrial large models according to the present invention have the following beneficial effects:
[0042] The present invention preprocesses the target corpus to obtain a standard corpus, performs scenario analysis on the standard corpus to obtain the tag information corresponding to the standard corpus, and then calculates the standard directory value corresponding to the target corpus according to the tag information. Since the standard directory value is calculated based on the tag information of the standard corpus and is unique, the corresponding standard directory can be generated according to the standard directory value later. The target corpus can be directly stored in the standard directory, or the corresponding corpus can be quickly queried according to the standard directory, without having to search sequentially according to the classification directory of the corpus, effectively improving the storage and call efficiency of the corpus and being able to improve the usage efficiency of the corpus. Description of the Drawings
[0043] Figure 1 It shows a flowchart of the corpus directory management method for industrial large models according to the present invention.
[0044] Figure 2 It shows a structural block diagram of the corpus directory management system for industrial large models according to the present invention. Detailed Embodiments
[0045] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0046] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0047] The corpus directory management method and system for industrial large models of the present invention preprocess the target corpus to obtain a standard corpus, perform scenario analysis on the standard corpus to obtain the corresponding tag information of the standard corpus, and then calculate the standard directory value corresponding to the target corpus according to the tag information. Since the standard directory value is calculated based on the tag information of the standard corpus and is unique, the corresponding standard directory can be generated according to the standard directory value later. The target corpus can be directly stored in the standard directory, or the corresponding corpus can be quickly queried according to the standard directory, without having to search sequentially according to the classification directory of the corpus, effectively improving the storage and invocation efficiency of the corpus and being able to improve the usage efficiency of the corpus.
[0048] A computer program is stored on the storage medium of the present invention, and when the computer program is executed by a processor, the following corpus management method is implemented. The storage medium includes: various media such as read-only memory (ROM), random access memory (RAM), magnetic disk, USB flash drive, memory card, or optical disc that can store program code.
[0049] Any combination of one or more storage media can be adopted. The storage medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0050] A computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take many forms, including - but not limited to - electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0051] The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including - but not limited to - wireless, wireline, optical fiber cable, RF, and the like, or any suitable combination of the foregoing.
[0052] The computer program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, Smalltalk, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0053] The present invention will be described below with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when the computer program instructions are executed by the processor of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams.
[0054] These computer program instructions can also be stored in a computer-readable medium that causes a computer, other programmable data processing apparatus, or other device to operate in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0055] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide a process for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0056] The terminal of the present invention includes a processor and a memory.
[0057] The memory is used for storing computer programs; preferably, the memory includes various media such as ROM, RAM, magnetic disks, USB flash drives, memory cards, or optical discs that can store program codes.
[0058] The processor is connected to the memory and is used for executing the computer programs stored in the memory, so that the terminal executes the following method for managing a corpus directory for an industrial large model.
[0059] Preferably, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0060] As Figure 1 shown, in an embodiment, the present invention discloses a method for managing a corpus directory for an industrial large model, including the following steps:
[0061] S100. Obtain the basic information of the target corpus, and preprocess the target corpus according to the basic information to obtain a standard corpus.
[0062] In this embodiment, when it is necessary to store the target corpus, the basic information of the target corpus is first obtained to facilitate subsequent preprocessing of the target corpus according to the basic information and ensure the final storage effect.
[0063] In some embodiments, the preprocessing the target corpus according to the basic information to obtain a standard corpus includes:
[0064] Performing a field scan on the target corpus to remove invalid fields and noise fields in the target corpus to obtain a first corpus;
[0065] Obtaining a sensitive word library corresponding to the target corpus according to the basic information, comparing the first corpus with the sensitive word library, and replacing sensitive words existing in the first corpus with tag words to obtain a corresponding standard corpus;
[0066] Wherein, each of the tag words corresponds to one of the sensitive words.
[0067] After determining the target corpus, first perform a field scan on the target corpus to remove invalid fields and noise fields therein to obtain a first corpus, avoiding the adverse effects of invalid fields and noise fields on subsequent processing. Among them, invalid fields include blank fields and duplicate fields, and noise fields include common noise information, such as advertisements, promotion links, special symbols, etc. Then, perform a comparison and matching in the sensitive word library according to the basic information of the target corpus to replace sensitive words with compliant tag words to obtain a corresponding standard corpus, avoiding the appearance of sensitive words in the target corpus during storage and affecting the use of the corpus.
[0068] It should be noted that after obtaining the standard corpus, it further includes converting the standard corpus into a standard format according to the standard encoding format to ensure that the text in the standard corpus has a unified format, facilitating subsequent processing of the standard corpus and ensuring text consistency.
[0069] S200. Performing a scenario analysis on the standard corpus to obtain corresponding tag information.
[0070] In still other embodiments, the basic information at least includes industry information, modality information, and language information. The performing a scenario analysis on the standard corpus to obtain corresponding tag information includes:
[0071] Obtaining the industry type of the target corpus according to the industry information, converting the industry type into a preset encoding format, and generating a first tag matrix corresponding to the standard corpus;
[0072] Obtain the modality type in the target corpus according to the modality information, convert the modality type into a preset encoding format, and generate a second label matrix corresponding to the standard corpus;
[0073] Obtain the language type of the target corpus according to the language information, convert the language type into a preset encoding format, and generate a third label matrix corresponding to the standard corpus;
[0074] Integrate the first label matrix, the second label matrix, and the third label matrix in sequence into a label matrix to obtain the label information;
[0075] Among them, the number of rows and columns of the first label matrix, the second label matrix, and the third label matrix are the same.
[0076] In this embodiment, since the standard corpus generally includes industry information, modality information, and language information, the industry information is the industry in which the standard corpus is used. For example, the corpus in the financial industry may contain information such as loans, investments, and financial products; the corpus in the medical industry may contain information such as diseases, drugs, and treatment methods. The modality information is organized according to the type and usage scenario of the corpus data, including text modality, image modality, and audio modality. The text modality includes text information such as articles, news, and blogs; the image modality includes image information such as photos, images, and videos; the audio modality includes audio information such as audio, phone recordings, and podcasts. The language information organizes the corpus according to the language type and region. Corpus data in different languages such as English, Chinese, and Spanish can be organized in different language directories respectively.
[0077] After determining the basic information of the target corpus, obtain the industry type of the target corpus according to the industry information, convert the industry type into a preset encoding format, and generate a first label matrix corresponding to the standard corpus. Among them, the industry type is the industry name corresponding to the target corpus. After converting the industry type into the corresponding preset encoding format, an industry code is obtained. Then, according to the preset matrix size, the industry code is used to establish the first label matrix to facilitate subsequent calculation and processing of the first label matrix.
[0078] Similarly, after determining the corresponding language modality information and language information according to the basic information of the target corpus, convert the modality type and the language type into the corresponding modality code and language code in the preset encoding format, and convert them into the corresponding second label matrix and third label matrix respectively.
[0079] It should be noted that the sizes of the first, second, and third label matrices are determined according to preset rules or the size of the target corpus, primarily to ensure uniformity in subsequent calculations. If there are insufficient codes, blank codes are added to ensure consistency in the resulting matrix specifications and facilitate subsequent calculations.
[0080] S300: Calculate the catalog value of the standard corpus according to the tag information to obtain a standard catalog value of the standard corpus.
[0081] In some other embodiments, calculating the catalog value of the standard corpus according to the tag information to obtain the standard catalog value of the standard corpus includes:
[0082] Calculating the same-position average of the first label matrix, the second label matrix, and the third label matrix to obtain an average label matrix;
[0083] Obtaining a corpus sequence of the standard corpus and generating a corpus index value of the corpus sequence;
[0084] Calculating the ratios of the eigenvalues corresponding to the first label matrix, the second label matrix, the third label matrix, and the average label matrix to the corpus index value, respectively, to obtain a first directory value, a second directory value, a third directory value, and an average directory value;
[0085] combining the first catalog value, the second catalog value, the third catalog value, and the average catalog value to generate a catalog sequence;
[0086] The directory index value of the directory sequence is calculated, the standard directory value of the standard corpus is determined according to the corpus index value and the directory index value, and a mapping relationship between the standard directory value and the standard corpus is established.
[0087] In this embodiment, after the first label matrix A, the second label matrix B, and the third label matrix C are calculated, the elements of the corresponding rows and columns of the first label matrix A, the second label matrix B, and the third label matrix C are first accumulated and averaged to obtain the average label matrix D. The specific formula is as follows:
[0088] , where i represents the i-th row of the matrix and j represents the j-th column of the matrix.
[0089] After calculating the average label matrix, obtain the corpus sequence of the standard corpus and generate the corpus index value of the corpus sequence. Specifically, first encode the standard corpus according to a preset format, and then arrange it in the encoding order to obtain the encoded corpus sequence. Then calculate the corpus index value through the corpus sequence, such as the sequence mean. After that, calculate the eigenvalues corresponding to the first label matrix, the second label matrix, the third label matrix, and the average label matrix respectively to obtain the first eigenvalue, the second eigenvalue, the third eigenvalue, and the average eigenvalue. Calculate the ratios of the first eigenvalue, the second eigenvalue, the third eigenvalue, and the average eigenvalue to the corpus index value respectively to obtain the first catalog value, the second catalog value, the third catalog value, and the average catalog value. Then combine the first catalog value, the second catalog value, the third catalog value, and the average catalog value together to form a catalog sequence. Since the first label matrix, the second label matrix, the third label matrix, the average label matrix, and the corpus sequence corresponding to different target corpora are all different, calculate the ratios between the eigenvalues of the first label matrix, the second label matrix, the third label matrix, the average label matrix and the corpus index value of the corpus sequence to obtain the first catalog value, the second catalog value, the third catalog value, and the average catalog value respectively. Combining the first catalog value, the second catalog value, the third catalog value, and the average catalog value together can obtain a catalog sequence with unique attributes for the target corpus, which is convenient for quickly locating the catalog of the target corpus.
[0090] Further, the generating the catalog sequence by combining the first catalog value, the second catalog value, the third catalog value, and the average catalog value includes:
[0091] Insert delimiters at both the beginning and the end of the first catalog value, the second catalog value, the third catalog value, and the average catalog value, and then arrange them in order to generate the catalog sequence.
[0092] After obtaining the catalog sequence, calculate the catalog index value of the catalog sequence, so as to determine the standard catalog value of the standard corpus according to the corpus index value and the catalog index value, and establish a mapping relationship between the standard catalog value and the standard corpus.
[0093] In this embodiment, since the catalog index value is the unique identifier of the catalog sequence, in order to further determine the unique catalog, determine the standard catalog value uniquely corresponding to the standard corpus according to the corpus index value of the standard corpus and the catalog index value of the catalog sequence. After establishing a mapping relationship between the standard catalog value and the standard corpus, the standard catalog value uniquely corresponding to the standard corpus can be determined, which is convenient for generating the corresponding target catalog according to the standard catalog value and quickly querying the corresponding standard corpus according to the standard catalog value, realizing the quick query and storage of the standard corpus.
[0094] Further, the standard directory value of the standard corpus is the sequence combination value of the corpus index value and the directory index value. For example, if the corpus index value is X and the directory index value is Y, then the standard directory value is (X, Y), which further ensures the uniqueness of the standard directory value.
[0095] Among them, both the corpus index value and the directory index value are sequences, thus ensuring that the standard directory value will not be repeated.
[0096] S400. Generate a corresponding target directory according to the standard directory value, and store the standard corpus after establishing a mapping with the target directory.
[0097] In some embodiments, the step of generating a corresponding target directory according to the standard directory value and storing the standard corpus after establishing a mapping with the target directory includes:
[0098] Obtain the corresponding directory path in the directory library according to the standard directory value, and generate a target directory corresponding to the standard corpus according to the directory path;
[0099] After establishing a mapping between the standard corpus and the target directory, store the standard corpus in the target directory.
[0100] After determining the standard directory value of the standard corpus, obtain the corresponding directory path in the directory library according to the standard directory value, so as to generate a target directory corresponding to the standard corpus according to the directory path. Then, after establishing a mapping between the standard corpus and the target directory, directly store the standard corpus in the target directory, thereby completing the rapid storage of the standard corpus. And when searching is needed, only by determining the standard directory value of the corpus can the corresponding directory path be determined, and the required corpus can be quickly found according to the directory path. There is no need to search and save in sequence according to the classification order of the corpus, effectively improving the storage and search efficiency of the corpus.
[0101] In some embodiments, the step of obtaining the corresponding directory path in the directory library according to the standard directory value and generating a target directory corresponding to the standard corpus according to the directory path includes:
[0102] Split the standard directory value into multiple sub-directory values with the corresponding quantity, and generate sub-paths in sequence for the multiple sub-directory values;
[0103] Combine the multiple sub-paths together to obtain the directory path, and generate a corresponding target directory according to the directory path.
[0104] To reduce complexity, the standard directory value is split into a preset number of sub-directory values in sequence, and then multiple sub-directory values are used to generate multiple sub-paths at the same time in order, and the multiple sub-paths are combined together in the original order to obtain a directory path. Then, a corresponding target directory can be generated according to the directory path to improve the generation efficiency of the target directory.
[0105] It should be noted that the protection scope of the corpus directory management method for industrial large models described in the present invention is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or reducing steps of the prior art and replacing steps according to the principles of the present invention is included in the protection scope of the present invention.
[0106] Such as Figure 2 shown, in an embodiment, the present invention provides a corpus directory management system for industrial large models, including:
[0107] A preprocessing module 201, configured to obtain basic information of a target corpus, and preprocess the target corpus according to the basic information to obtain a standard corpus;
[0108] An analysis module 202, configured to perform scenario analysis on the standard corpus to obtain corresponding tag information;
[0109] A calculation module 203, configured to calculate a directory value of the standard corpus according to the tag information to obtain a standard directory value of the standard corpus;
[0110] A directory generation module 204, configured to generate a corresponding target directory according to the standard directory value, and store the standard corpus after establishing a mapping with the target directory.
[0111] It should be noted that the structure and principle of the above corpus directory management system for industrial large models correspond one by one to the steps in the above corpus directory management method for industrial large models, so they will not be elaborated here.
[0112] It should be noted that the division of each module of the above system is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the x module can be a separately established processing element, or can be integrated in a certain chip of the above system. In addition, it can also be stored in the memory of the above system in the form of program code, and the function of the above x module can be called and executed by a certain processing element of the above system. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit of the hardware in the processor element or the instruction in the form of software.
[0113] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0114] It should be noted that the corpus directory management system for industrial large models of the present invention can implement the corpus directory management method for industrial large models of the present invention, but the implementation device of the corpus directory management method for industrial large models of the present invention includes but is not limited to the structure of the corpus directory management system for industrial large models listed in this embodiment. Any structural deformation and replacement of the prior art made according to the principles of the present invention are included in the protection scope of the present invention.
[0115] In summary, the corpus directory management method and system for industrial large models of the present invention preprocess the target corpus to obtain a standard corpus, perform scenario analysis on the standard corpus to obtain the corresponding tag information of the standard corpus, and then calculate the standard directory value corresponding to the target corpus according to the tag information. Since the standard directory value is calculated based on the tag information of the standard corpus and is unique, the corresponding standard directory can be generated according to the standard directory value later. The target corpus can be directly stored in the standard directory, or the corresponding corpus can be quickly queried according to the standard directory, without having to search sequentially according to the classification directory of the corpus, effectively improving the storage and call efficiency of the corpus and being able to improve the usage efficiency of the corpus. Therefore, the present invention effectively overcomes various shortcomings in the prior art and has high industrial utilization value.
[0116] The above embodiments are only illustrative of the principles and effects of the present invention, and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A method for managing a corpus directory for an industrial large model, characterized in that, It includes the following steps: Obtain the basic information of the target corpus, and preprocess the target corpus according to the basic information to obtain a standard corpus; Conduct a scenario analysis on the standard corpus to obtain corresponding label information; Calculate the catalog value of the standard corpus according to the label information to obtain the standard catalog value of the standard corpus; Generate a corresponding target catalog according to the standard catalog value, establish a mapping between the standard corpus and the target catalog, and then store them; The preprocessing of the target corpus according to the basic information to obtain a standard corpus includes: Conduct a field scan on the target corpus to remove invalid fields and noise fields in the target corpus to obtain a first corpus; Obtain the sensitive word library corresponding to the target corpus according to the basic information, compare the first corpus with the sensitive word library, and replace the sensitive words existing in the first corpus with label words to obtain the corresponding standard corpus; Wherein, each label word corresponds to each sensitive word one by one; The basic information includes at least industry information, modality information and language information. The scenario analysis of the standard corpus to obtain corresponding label information includes: Obtain the industry type of the target corpus according to the industry information, convert the industry type into a preset coding format, and generate a first label matrix corresponding to the standard corpus; Obtain the modality type in the target corpus according to the modality information, convert the modality type into a preset coding format, and generate a second label matrix corresponding to the standard corpus; Obtain the language type of the target corpus according to the language information, convert the language type into a preset coding format, and generate a third label matrix corresponding to the standard corpus; Integrate the first label matrix, the second label matrix and the third label matrix in sequence into a label matrix to obtain the label information; Wherein, the number of rows and columns of the first label matrix, the second label matrix and the third label matrix are the same; The calculation of the catalog value of the standard corpus according to the label information to obtain the standard catalog value of the standard corpus includes: Calculate the co-location average value of the first label matrix, the second label matrix and the third label matrix to obtain an average label matrix; Obtain the corpus sequence of the standard corpus and generate the corpus index value of the corpus sequence; Calculate the ratios between the eigenvalues corresponding to the first label matrix, the second label matrix, the third label matrix and the average label matrix and the corpus index value respectively to obtain a first catalog value, a second catalog value, a third catalog value and an average catalog value respectively; Combine the first catalog value, the second catalog value, the third catalog value and the average catalog value together to generate a catalog sequence; Calculate the catalog index value of the catalog sequence, determine the standard catalog value of the standard corpus according to the corpus index value and the catalog index value, and establish a mapping relationship between the standard catalog value and the standard corpus.
2. The corpus directory management method for industrial large models according to claim 1, wherein Generating a corresponding target directory according to the standard directory value, and storing the standard corpus after establishing a mapping with the target directory, including: Correspondingly obtaining a corresponding directory path in the directory library according to the standard directory value, and generating a target directory corresponding to the standard corpus according to the directory path; After establishing a mapping between the standard corpus and the target directory, storing the standard corpus in the target directory.
3. The corpus directory management method for industrial large models according to claim 2, wherein, The correspondingly obtaining a corresponding directory path in the directory library according to the standard directory value, and generating a target directory corresponding to the standard corpus according to the directory path, including: Splitting the standard directory value into a plurality of sub-directory values corresponding in number, and generating sub-paths in sequence from the plurality of sub-directory values; Combining the plurality of sub-paths together to obtain the directory path, and generating a corresponding target directory according to the directory path.
4. The method for managing a corpus directory for an industrial large model according to claim 2, wherein The combining the first directory value, the second directory value, the third directory value and the average directory value together to generate a directory sequence, including: Inserting delimiters at both the beginning and the end of the first directory value, the second directory value, the third directory value and the average directory value, and then arranging them in order to generate the directory sequence.
5. The method for corpus directory management for industrial large models according to claim 2, wherein, The standard directory value of the standard corpus is the sequence combination value of the corpus index value and the directory index value.
6. A corpus directory management system for industrial large models, characterized in that, Including: A preprocessing module, configured to obtain basic information of a target corpus, and preprocess the target corpus according to the basic information to obtain a standard corpus; An analysis module, configured to perform scenario analysis on the standard corpus to obtain corresponding label information; A calculation module, configured to calculate a directory value for the standard corpus according to the label information to obtain the standard directory value of the standard corpus; A directory generation module, configured to generate a corresponding target directory according to the standard directory value, and store the standard corpus after establishing a mapping with the target directory; Wherein, the preprocessing the target corpus according to the basic information to obtain a standard corpus includes: Performing a field scan on the target corpus to remove invalid fields and noise fields in the target corpus to obtain a first corpus; Obtaining a sensitive word library corresponding to the target corpus according to the basic information, comparing the first corpus with the sensitive word library, and replacing sensitive words existing in the first corpus with label words to obtain a corresponding standard corpus; Wherein, each label word corresponds to one of the sensitive words; The basic information at least includes industry information, modality information and language information. The performing scenario analysis on the standard corpus to obtain corresponding label information includes: Obtaining the industry type of the target corpus according to the industry information, converting the industry type into a preset coding format, and generating a first label matrix corresponding to the standard corpus; Obtaining the modality type in the target corpus according to the modality information, converting the modality type into a preset coding format, and generating a second label matrix corresponding to the standard corpus; Acquiring a language type of the target corpus according to the language information, converting the language type into a preset encoding format, and generating a third label matrix corresponding to the standard corpus; Integrating the first label matrix, the second label matrix, and the third label matrix in sequence into a label matrix to obtain the label information; The rows and columns of the first label matrix, the second label matrix and the third label matrix are the same; The calculating the catalog value of the standard corpus according to the tag information to obtain the standard catalog value of the standard corpus includes: Calculating the same-position average of the first label matrix, the second label matrix, and the third label matrix to obtain an average label matrix; Obtaining a corpus sequence of the standard corpus and generating a corpus index value of the corpus sequence; Calculating the ratios of the eigenvalues corresponding to the first label matrix, the second label matrix, the third label matrix, and the average label matrix to the corpus index value, respectively, to obtain a first directory value, a second directory value, a third directory value, and an average directory value; combining the first catalog value, the second catalog value, the third catalog value, and the average catalog value to generate a catalog sequence; The directory index value of the directory sequence is calculated, the standard directory value of the standard corpus is determined according to the corpus index value and the directory index value, and a mapping relationship between the standard directory value and the standard corpus is established.
Citation Information
Patent Citations
Book management method and device
CN116244424A
Corpus management method, electronic equipment, storage medium and program product
CN119226289A