Corpus operation platform of industrial large model
By designing a corpus operation platform for industrial large models, the problem of low corpus management efficiency in the existing technology is solved, and more efficient corpus use and better user experience is achieved.
Patent Information
- Application Number
- CN202510595395.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing corpus management platform cannot effectively manage different users and corpus, resulting in poor user experience and low corpus use efficiency.
Design a corpus operation platform for industrial large-scale models, including user management module, corpus management module, corpus traceability module and corpus production module. The platform performs unified user management through the user management module. The corpus management module adaptively generates corpus catalogs and tool catalogs based on user information. The corpus traceability module provides service provider information. The corpus production module generates target corpus according to user needs.
It improves the efficiency of the use of corpus, enhances the user experience, and meets users' multiple needs for corpus.
Smart Images

Figure CN120146159A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large models, and particularly relates to a corpus operation platform for industrial large models. Background Art
[0002] Artificial intelligence large models refer to a class of artificial intelligence models with a large number of parameters constructed by artificial neural networks. They are usually pre-trained on a large amount of data through self-supervised learning or semi-supervised learning, and then further optimize their performance and capabilities through methods such as instruction tuning and human alignment. Large models have the characteristics of a large number of parameters, large training data, and large computing resources, and have the abilities to solve general tasks, follow human instructions, and perform complex reasoning. The main categories of artificial intelligence large models include: large language models, vision large models, multimodal large models, and basic science large models, etc. At present, large models have been widely applied in multiple fields, including search engines, intelligent agents, related vertical industries, and basic science, etc., promoting the intelligent development of various industries.
[0003] And corpus is an important material for training artificial intelligence large models. Generally, corpus refers to the instances and data sets used in linguistic research and natural language processing. It can be written text, spoken records, or other structured data, and is usually used for analyzing language phenomena, supporting machine translation, speech recognition, automatic text summarization, and other tasks. Usually, a corpus is a set of text or speech data that has been collected, sorted, and annotated. These data are used in linguistic research to analyze language usage rules, vocabulary changes, and grammatical structures, etc. In natural language processing, corpus is the basic data source for training and testing models, supporting functions such as machine translation, speech recognition, and sentiment analysis.
[0004] Due to the rapid development of current large models, the corresponding demand for corpus has also increased. However, most of the existing corpus management platforms only manage the corpus uniformly, and cannot effectively manage different users and corpus according to user needs, which affects the user experience and also reduces the usage efficiency of the corpus. Summary of the Invention
[0005] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a corpus operation platform for industrial large models, which is used to solve the problems of low corpus management efficiency in the prior art, resulting in poor user experience and affecting the usage efficiency of the corpus.
[0006] To achieve the above purpose and other related purposes, the present invention provides a corpus operation platform for industrial large models, including: A user management module, which is used for user registration and uniformly manages the registered users; A corpus management module, which is used to provide various directory information for the users registered in the user management module for the users to view and manage the corpus; A corpus traceability module, which is communicatively connected to the corpus management module and is used to provide various service provider information for the users; A corpus production module, which is used to generate corresponding target corpus according to the corpus plan formulated by the users; Wherein, the corpus management module is further used to adaptively generate corresponding corpus directories and tool directories according to the user information.
[0007] In an embodiment of the present invention, the user management module includes a user registration unit, a user login unit, a user authentication unit, a permission management unit and a log management unit. The user registration unit is used for new user registration. The user login unit is communicatively connected to the user authentication unit and is used to authenticate the user when the user logs in, determine the target permission information of the user according to the authentication result, and allocate corresponding system menus to the user according to the target permission information. The permission management unit is communicatively connected to the user authentication unit and is used to allocate corresponding target permissions to the user according to the target permission information after the user completes registration. The log management unit is communicatively connected to the user login unit and is used to record the operation situation of the user and generate a record log after the user logs in.
[0008] In an embodiment of the present invention, the registration process of the user in the user registration unit includes: After the user inputs registration information according to the registration interface, the user is activated and verified according to the registration information; After the verification is passed, the basic information of the user is collected, and the basic information is parsed for permissions to obtain corresponding target permission information; The target permission information is sent to the permission management unit, and the permission management unit allocates corresponding target permissions to the registered user.
[0009] In an embodiment of the present invention, the parsing of the basic information for permissions to obtain corresponding target permission information includes: The basic information is converted into a standard format to obtain standard information, and the standard information is input into the database to update the user information table in the database. The user information table is used to store the processed standard information of each user; Calculate the weight coefficient of each data index according to the updated user information table; Calculate the target permission information of the currently registered user according to the weight coefficient and the standard information.
[0010] In an embodiment of the present invention, calculating the weight coefficient of each data index according to the updated user information table includes: Calculating the mean and standard deviation of the data index corresponding to each user in the user information table, where the data index of the user includes at least one of corpus call information, corpus download information, corpus usage frequency, and recharge information; Performing a first normalization process on the index information according to the mean and the standard deviation to obtain a first standard user information table; Constructing a covariance matrix according to the data in the first standard user information table; Performing eigenvalue decomposition on the covariance matrix to obtain an eigenvalue matrix and an eigenvector matrix; Calculating the load coefficient corresponding to the data index of each user according to the eigenvector matrix; Calculating the first weight of each index information according to the load coefficient and the eigenvalue matrix; Performing a second normalization process on the data index in the user information table to obtain a second standard user information table; Calculating the proportion information of each user in the second standard user information table on different data indexes; Obtaining the difference coefficient corresponding to each data index according to the proportion information, and calculating the second weight corresponding to each data index according to the difference coefficient; Calculating the weight coefficient corresponding to the data index according to the first weight and the second weight.
[0011] In an embodiment of the present invention, the permission management unit is further configured to perform a security specification verification on the user at random times during the user's use process, and adjust the user's permissions according to the security specification verification result; Wherein, when the number of times the security specification verification result of the user appears incorrect reaches a first preset number of times, the user's permission level is correspondingly reduced; When the number of consecutive correct times of the security specification verification result of the user reaches a second preset number of times, the user's permission level is increased.
[0012] In an embodiment of the present invention, the corpus management module includes a corpus directory unit, a tool directory unit, and a service provider directory unit. The corpus directory unit is used to generate a corpus directory corresponding to the target permission information of the user. The tool directory unit is communicatively connected to the corpus directory unit and is used to generate a matching tool directory according to the corpus directory. The tool directory is used for the user to perform different processes on the corpus directory. The service provider directory is used to provide the service provider information corresponding to each corpus according to the corpus directory.
[0013] In an embodiment of the present invention, the process by which the tool directory unit generates a matching tool directory according to the corpus directory includes: Obtain the first usage information of the tools corresponding to each corpus in the corpus directory, and calculate the first sorting information of the matching tools corresponding to each corpus according to the first usage information. The matching tool is the recommended tool for the corpus. Obtain the second usage information of the remaining tools in the existing toolbox, and calculate the second sorting information of each of the tools according to the second usage information. Add the first sorting information corresponding to each corpus to the second sorting information respectively and re-sort to generate target sorting information, and generate a tool directory corresponding to each corpus according to the result of the target sorting information.
[0014] In an embodiment of the present invention, the obtaining the first usage information of the tools corresponding to each corpus in the corpus directory and calculating the first sorting information of the matching tools corresponding to each corpus according to the first usage information includes: Obtain the first usage frequency, first evaluation score, first update frequency, first accuracy, and first single-use duration of the matching tools corresponding to each corpus in the corpus directory, and convert the first usage frequency, the first evaluation score, the first update frequency, the first accuracy, and the first single-use duration information into first standard information. Perform weighted calculation on the first standard information according to a first preset weight value to obtain the first sorting information of the matching tools corresponding to each corpus. The obtaining the second usage information of the remaining tools in the existing toolbox and calculating the second sorting information of each of the tools according to the second usage information includes: Obtain the second usage frequency, second evaluation information, second update frequency, second accuracy, and second single-use duration of the remaining tools in the existing toolbox, and convert the second usage frequency, the second evaluation information, the second update frequency, the second accuracy, and the second single-use duration into second standard information. Perform weighted calculation on the second standard information according to the second preset weight value to obtain the second sorting information of each tool; After adding the first sorting information corresponding to each corpus into the second sorting information respectively and re - sorting to generate the target sorting information, and generating a tool catalog corresponding to each corpus according to the result of the target sorting information, including: Respectively add the first sorting information of the tools corresponding to each corpus after the second sorting information and re - sort, and obtain the target sorting information corresponding to each corpus according to the sorting result; Generate a tool catalog corresponding to each corpus according to the target sorting information.
[0015] In an embodiment of the present invention, it is used to provide detailed information of the service provider corresponding to the corpus, sort the detailed information of the service provider according to the sorting result of the corpus, establish a one - to - one mapping relationship between the corpus and the service provider information, and the corpus production module is also communicatively connected to the corpus traceability module for generating the corresponding target corpus according to the provided service provider information.
[0016] As described above, the corpus operation platform of the industrial large - model of the present invention has the following beneficial effects: Compared with the prior art, the present invention enables users to directly register through the user management module and uniformly manage users. After the user completes registration, the corpus management module provides corresponding corpus catalogs and tool catalogs for different users according to the user's information, so as to facilitate different users to quickly query the adapted corpus catalogs and tool catalogs, facilitate users to call different corpora, effectively improve the usage efficiency of the corpora, enhance the user experience. At the same time, the corpus traceability module can provide corresponding service provider information for users according to the corpus information, and the corpus production module specifies the corresponding corpus plan according to the user's needs and generates the corresponding target corpus, meeting the multiple needs of users for corpora from different aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It shows a structural block diagram of the corpus operation platform of the industrial large - model of the present invention in an embodiment.
[0018] Figure 2 It shows a structural block diagram of the user management module in the corpus operation platform of the industrial large - model of the present invention.
[0019] Figure 3 It shows a flowchart of the user registration process in the corpus operation platform of the industrial large - model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following describes the implementation manners of the present invention through specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0021] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. The diagrams only show the components related to the present invention, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0022] For the corpus operation platform of the industrial large model of the present invention, the present invention enables users to directly register through the user management module and uniformly manage users. After the user completes registration, the corpus management module provides corresponding corpus directories and tool directories for different users according to the user information, so as to facilitate different users to quickly query the suitable corpus directories and tool directories, facilitate the users to call different corpora, effectively improve the usage efficiency of the corpora, enhance the user experience. At the same time, the corpus traceability module can provide corresponding service provider information for the users according to the corpus information, and the corpus production module specifies corresponding corpus plans according to the user requirements and generates corresponding target corpora, meeting the multiple requirements of users for corpora from different aspects.
[0023] As Figure 1 shown, in one embodiment, the present invention provides a corpus operation platform for an industrial large model, including: A user management module 1, which is used for user registration and uniformly manages the registered users; A corpus management module 2, which is used to provide various directory information for the users registered in the user management module 1 for the users to view and manage the corpora; A corpus traceability module 3, which is communicatively connected to the corpus management module 2 and is used to provide various service provider information for the users; A corpus production module 4, which is used to generate corresponding target corpora according to the corpus plans formulated by the users; Wherein, the corpus management module is further used to adaptively generate corresponding corpus directories and tool directories according to the user information.
[0024] In this embodiment, the user management module facilitates direct user registration and unified user management. After the user completes registration, the corpus management module 2 provides corresponding corpus directories and tool directories for different users according to the user information, so as to facilitate different users to quickly query the adapted corpus directories and tool directories, facilitate the user's call of different corpora, effectively improve the usage efficiency of the corpora, enhance the user experience. At the same time, the corpus traceability module can provide corresponding service provider information for the user according to the corpus information, and the corpus production module specifies the corresponding corpus plan and generates the corresponding target corpus according to the user's needs, meeting the user's multiple needs for corpora from different aspects.
[0025] In some embodiments, referring to Figure 2 , the user management module 1 includes a user registration unit 11, a user login unit 12, a user authentication unit 13, a permission management unit 14, and a log management unit 15. The user registration unit 11 is used for new user registration. The user login unit 12 is communicatively connected to the user authentication unit 13, and is used to authenticate the user when the user logs in, determine the target permission information of the user according to the authentication result, and allocate the corresponding system menu for the user according to the target permission information. The permission management unit 14 is communicatively connected to the user authentication unit 13, and is used to allocate the corresponding target permission to the user according to the target permission information after the user completes registration. The log management unit 15 is communicatively connected to the user login unit 12, and is used to record the operation situation of the user and generate a record log after the user logs in.
[0026] In this embodiment, the user registration unit 11 is used for the user to complete the registration process on the entire platform. After completion of registration, when the user logs in to the platform through the user login unit 12, the user authentication unit 13 authenticates the user information, and determines the target permission information of the user according to the authentication result after the authentication is completed. Then, the permission management unit 14 allocates the corresponding target permission for the user according to the target permission information. At the same time, during the user's use on the user platform, the log management unit 15 records the operation situation of the user and generates the corresponding record log, so as to manage the user's behavior and facilitate timely traceability in case of problems.
[0027] Further, referring to Figure 3 , the registration process of the user in the user registration unit 11 includes the following steps: S301. After the user inputs registration information according to the registration interface, activate and verify the user according to the registration information; S302. After the verification passes, collect the basic information of the user, and perform permission parsing on the basic information to obtain the corresponding target permission information; S303. Send the target permission information to the permission management unit, and allocate the corresponding target permission to the user who has completed registration through the permission management unit.
[0028] In this embodiment, the user registration unit 11 provides a registration interface for user registration. The registration interface provides an input for the user's registration information, which includes registration information such as user name, user password, random verification code, and email verification. At the same time, the basic information of the user is collected synchronously on the registration interface. The basic information includes corpus call information, corpus download information, corpus usage frequency, and recharge information, so as to facilitate subsequent permission parsing of the user based on the basic information to obtain the corresponding target permission information of the user. Then, the target permission information is sent to the permission management unit 14, so as to allocate the corresponding target permission to the user who has completed registration, facilitate the user to quickly obtain the corresponding permission after logging in, and ensure that the user can use the corpus on the platform within the scope of the permission.
[0029] In some embodiments, the performing permission parsing on the basic information to obtain the corresponding target permission information includes: Convert the basic information into a standard format to obtain standard information, and input the standard information into the database to update the user information table in the database. The user information table is used to store the processed standard information of each user; Calculate the weight coefficient of each data index according to the updated user information table; Calculate the target permission information of the currently registered user according to the weight coefficient and the standard information; Wherein, the basic information includes multiple data indexes.
[0030] In this embodiment, after obtaining the basic information of the user, in order to facilitate the calculation of the target weight information, first convert the basic information into standard information in a standard format to facilitate unified calculation of the basic information. Then, input the standard information into the database and update the user information table in the database. Then, calculate the weight coefficient of each data index in the basic information of the user according to the updated user information table, so as to facilitate subsequent calculation of the target permission information corresponding to the user according to the weight coefficient and each data index of the standard information.
[0031] Wherein, the conversion of the basic information of the user into a standard format is mainly to facilitate the subsequent standard calculation process. Specifically, it is achieved by using the means in the existing technology to standardize the basic information. This solution is not particularly limited here and will not be elaborated further.
[0032] In some other embodiments, calculating the weight coefficient of each data metric according to the updated user information table includes: Calculating the mean and standard deviation of the data metrics corresponding to each user in the user information table, where the data metrics of the user include at least one of corpus call information, corpus download information, corpus usage frequency, and recharge information; Performing a first standardization process on the metric information according to the mean and the standard deviation to obtain a first standard user information table; Constructing a covariance matrix based on the data in the first standard user information table; Performing eigenvalue decomposition on the covariance matrix to obtain an eigenvalue matrix and an eigenvector matrix; Calculating the load coefficient corresponding to the data metric of each user according to the eigenvector matrix; Calculating the first weight of each metric information according to the load coefficient and the eigenvalue matrix; Performing a second standardization process on the data metrics in the user information table to obtain a second standard user information table; Calculating the proportion information of each user in the second standard user information table on different data metrics; Obtaining the difference coefficient corresponding to each data metric according to the proportion information, and calculating the second weight corresponding to each data metric according to the difference coefficient; Calculating the weight coefficient corresponding to the data metric according to the first weight and the second weight.
[0033] In this embodiment, after performing standardization processing on the basic information of the user to obtain the standardized standard information, the standard information of the current user is added to the database to update the user information table, and then the weight coefficient of each data metric is calculated according to the updated user information table.
[0034] Specifically, first calculate the mean and standard deviation of the data metrics corresponding to each user in the updated user information table, and perform a first standardization process on the data metrics according to the mean and the standard deviation to obtain the updated first standard user information table.
[0035] Among them, the calculation process of the data metrics after the first standardization process satisfies the following formula: ; Among them, represents the value of the jth data metric of the ith user after the first standardization process, represents the value of the jth data metric of the ith user, represents the mean of the j-th data metric for all the said users, represents the standard deviation of the j-th data metric for all the said users.
[0036] After completing the first standardization process on the metric information, calculate the corresponding n×n covariance matrix based on the data metrics in the first standard user information table, where n represents the number of data metrics of the said users, to reflect the correlation between various data metrics. The elements in the covariance matrix satisfy the following formula: ; where, represents the covariance between the j-th data metric and the k-th data metric in the first standard user information table, represents the j-th data metric of the i-th user in the first standard user information table, represents the k-th data metric of the i-th user in the first standard user information table, and m represents the number of users in the first standard user information table.
[0037] Then perform eigenvalue decomposition on the covariance matrix of the components to obtain an eigenvalue array arranged from largest to smallest and the corresponding eigenvector array , and then calculate the loading coefficient of the j-th data metric according to the eigenvector array .
[0038] where, the relationship between the loading coefficient and the eigenvector array satisfies the following formula: ; where, represents the loading coefficient of the j-th data metric of the k-th said user.
[0039] After obtaining the loading coefficients of each data metric, calculate the first weight of the j-th data metric of the said users according to the loading coefficient and the eigenvalue array calculated above .
[0040] The calculation process of the first weight of the j-th data metric satisfies the following formula: ; where, represents the eigenvalue of the k-th user, represents the loading coefficient of the j-th data metric of the k-th user.
[0041] Furthermore, use the range transformation method to perform the second standardization process on the data metrics in the user information table to obtain the second standard user information table, which specifically satisfies the following formula: ; wherein, d ij represents the value of the j-th data index of the i-th user in the second standard user information table, represents the value of the j-th data index of the i-th user in the user information table, represents the maximum value of the j-th data index in the user information table, represents the minimum value of the j-th data index in the user information table.
[0042] After that, calculate the proportion information P ij of the j-th data index of the i-th user in the second standard user information table respectively, which specifically satisfies the following formula: ; m represents the number of users in the second standard user information table.
[0043] After that, calculate the coefficient of variation of each data index according to the proportion information, and the calculation process of the coefficient of variation satisfies the following formula: , ; wherein, y j represents the coefficient of variation of the j-th data index, and m represents the number of users in the second standard user information table.
[0044] After that, according to the coefficient of variation, the second weight of each index data in the second standard user information table can be obtained .
[0045] Its specific process satisfies the following formula: ; After calculating the first weight and the second weight of the data index of the user, further allocate the weight values to calculate the final weight coefficient, and the calculation process of the weight coefficient satisfies the following formula: ; wherein, represents the allocation coefficient of the first weight, represents the allocation coefficient of the second weight, and .
[0046] It should be noted that the allocation coefficient of the first weight and the allocation coefficient of the second weight can be preset or adjusted according to empirical values, which will not be elaborated here.
[0047] In some other embodiments, the permission management unit 14 is further configured to, during the user's usage process, verify the user's security specifications at random times, and adjust the user's permissions according to the results of the security specification verification; Wherein, when the number of times the security specification verification result of the user appears incorrect reaches a first preset number of times, the user's permission level is correspondingly reduced; When the number of consecutive correct times of the user's security specification verification result reaches a second preset number of times, the user's permission level is increased.
[0048] Among them, verifying the user's security specifications mainly refers to verifying the user's security specifications during the process of using the corpus on the platform to determine whether there are any operation security problems during the user's usage. If the number of operation security errors of the user reaches the first preset number of times, the user's permission level is correspondingly reduced to prevent the user's insecure operations from affecting the security of the corpus on the platform; if the number of consecutive correct times of the user's operation security reaches the second preset number of times, it is determined that the current user's operations have been in compliance with the specifications for a long time, and the user's permission level can be appropriately increased, thereby creating better conditions for users who use the platform in a standardized manner.
[0049] In still some other embodiments, the corpus management module 2 includes a corpus directory unit 21, a tool directory unit 22, and a service provider directory unit 23. The corpus directory unit 21 is configured to generate a corpus directory corresponding to the user's target permission information. The tool directory unit 22 is communicatively connected to the corpus directory unit 21 and is configured to generate a matching tool directory according to the corpus directory. The tool directory is used for the user to perform different processes on the corpus directory, and the service provider directory is configured to provide the service provider information corresponding to each corpus according to the corpus directory.
[0050] In this embodiment, after the user completes registration, the target permission information of the user is determined by the permission management unit 14. Then, the corpus directory unit 21 in the corpus management module 2 generates a corresponding and matching corpus directory according to the user's target permission information to determine that the user can obtain a corpus directory matching the permissions. At the same time, the tool directory unit 22 generates a tool directory corresponding to and matching the corpus directory, so as to facilitate the user to quickly call the tools in the tool directory to quickly process the corpus, which is convenient for the user to operate. The service provider directory generated by the service provider directory unit 23 facilitates the user to view the service provider information when calling the corpus and trace the source of the corpus, ensuring the security and compliance of the corpus.
[0051] In some embodiments, the process by which the tool directory unit generates a matching tool directory according to the corpus directory includes: Obtain the first usage information of the tools corresponding to each corpus in the corpus directory, and calculate the first sorting information of the matching tools corresponding to each corpus according to the first usage information, where the matching tool is the recommended tool for the corpus; Obtain the second usage information of the remaining tools in the existing toolbox, and calculate the second sorting information of each of the tools according to the second usage information; Add the first sorting information corresponding to each corpus to the second sorting information respectively and re - sort to generate the target sorting information, and generate the tool directory corresponding to each corpus according to the result of the target sorting information.
[0052] In this embodiment, after generating the corresponding corpus directory according to the user's target permission information, first obtain the first usage information of the tools corresponding to each corpus in the corpus directory, calculate the first sorting information of the matching tools corresponding to the corpus according to the first usage information, then obtain the second usage information of the remaining tools in the toolbox, and calculate the second sorting information of the remaining tools according to the second usage information. Then, by combining the first sorting information and the second sorting information of the matching tools corresponding to each corpus, the target sorting information corresponding to each corpus is obtained. According to the result of the target sorting information, the tool directory corresponding to each corpus can be generated, so as to ensure that each corpus has the most suitable tool directory, which is convenient for users to select appropriate tools for invocation.
[0053] In some embodiments, the obtaining the first usage information of the tools corresponding to each corpus in the corpus directory and calculating the first sorting information of the matching tools corresponding to each corpus according to the first usage information includes: Obtain the first usage frequency, the first evaluation score, the first update frequency, the first accuracy, and the first single - use duration of the matching tools corresponding to each corpus in the corpus directory, and convert the first usage frequency, the first evaluation score, the first update frequency, the first accuracy, and the first single - use duration information into first standard information; Perform weighted calculation on the first standard information according to the first preset weight value to obtain the first sorting information of the matching tools corresponding to each corpus; The obtaining the second usage information of the remaining tools in the existing toolbox and calculating the second sorting information of each of the tools according to the second usage information includes: Obtain the second usage frequency, the second evaluation information, the second update frequency, the second accuracy, and the second single - use duration of the remaining tools in the existing toolbox, and convert the second usage frequency, the second evaluation information, the second update frequency, the second accuracy, and the second single - use duration into second standard information; Performing weighted calculation on the second standard information according to the second preset weight value to obtain the second sorting information of each tool; Adding the first sorting information corresponding to each corpus to the second sorting information respectively, and then re - sorting to generate the target sorting information, and generating a tool catalog corresponding to each corpus according to the result of the target sorting information, including: Respectively adding the first sorting information of the matching tool corresponding to each corpus after the second sorting information and re - sorting, and obtaining the target sorting information corresponding to each corpus according to the sorting result; Generating a tool catalog corresponding to each corpus according to the target sorting information.
[0054] In this embodiment, after the user registration is completed and the corpus catalog corresponding to the user is generated, first obtain the matching tools corresponding to each corpus inside the corpus catalog, that is, the first usage frequency, the first evaluation score, the first update frequency, the first accuracy, and the first single - use duration of the recommended tools, and perform normalization processing to obtain the first standard information. Similarly, obtain the second usage frequency, the second evaluation information, the second update frequency, the second accuracy, and the second single - use duration of the remaining tools in the existing toolbox, where the remaining tools are the tools except the recommended tools, and convert the second usage frequency, the second evaluation information, the second update frequency, the second accuracy, and the second single - use duration into the second standard information.
[0055] Among them, both the first evaluation information and the second evaluation information are evaluated using score values to facilitate subsequent calculation of the sorting result. And the normalization calculation processes of the usage frequency, the evaluation score, the update frequency, the single - use duration, and the accuracy all satisfy the following formula: ; x i represents the i - th original data, X i represents the i - th standard data after normalization processing, x max represents the maximum data in the same category.
[0056] After obtaining the first standard information and the second standard information, perform weighted calculation on the first standard information according to the first preset weight value to obtain the first sorting information of the matching tool corresponding to each corpus. At the same time, perform weighted calculation on the second standard information according to the second preset weight value to obtain the second sorting information of each tool. Then, add the first sorting information of the matching tool corresponding to each corpus to the second sorting information respectively and re-sort, and obtain the target sorting information corresponding to each corpus according to the sorting result; generate the tool catalog corresponding to each corpus according to the target sorting information, so that the most suitable tool catalog corresponding to each corpus can be obtained, which is convenient for users to quickly call the tool to use the corpus and improves the user's usage efficiency of the corpus. It should be noted that the first preset weight value and the second preset weight value are preset empirical values, and can also be adjusted according to the experience value according to the change of data in the future.
[0057] In some embodiments, the corpus traceability module 3 is also communicatively connected to the corpus management module 2, and is used to provide the detailed information of the service provider corresponding to the corpus, sort the detailed information of the service provider according to the sorting result of the corpus, and establish a one-to-one mapping relationship between the corpus and the service provider information. And the corpus production module 4 is also communicatively connected to the corpus traceability module 3, and is used to generate the corresponding target corpus according to the provided service provider information. By sorting the service provider information according to the sorting result of the corpus, it is convenient for users to quickly find the corresponding service provider information when calling the corpus. At the same time, the corpus production module 4 is also convenient for generating the corpus according to the service provider information provided by the corpus traceability module 4 after receiving the corpus production plan specified by the user.
[0058] In summary, the corpus operation platform of the industrial large model of the present invention enables users to directly register through the user management module and uniformly manage users. After the user completes registration, the corpus management module provides corresponding corpus catalogs and tool catalogs for different users according to the user's information, so as to facilitate different users to quickly query the adapted corpus catalogs and tool catalogs, facilitate users to call different corpora, effectively improve the usage efficiency of the corpus, enhance the user experience. At the same time, the corpus traceability module can provide corresponding service provider information for users according to the corpus information, and the corpus production module specifies the corresponding corpus plan according to the user's needs and generates the corresponding target corpus, meeting the multiple needs of users for the corpus from different aspects; Therefore, the present invention effectively overcomes various shortcomings in the prior art and has high industrial utilization value.
[0059] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A corpus operation platform for industrial large models, characterized by: include: User management module, used for user registration and unified management of registered users; A corpus management module, used to provide various directory information for the user registered in the user management module, so that the user can view and manage corpus; A corpus tracing module, which is in communication with the corpus management module and is used to provide the user with various service provider information; A corpus production module, used to generate corresponding target corpus according to the corpus plan formulated by the user; The corpus management module is further used to adaptively generate a corresponding corpus catalog and tool catalog according to the user information.
2. The corpus operation platform for industrial large models according to claim 1 is characterized in that: The user management module includes a user registration unit, a user login unit, a user authentication unit, a permission management unit and a log management unit. The user registration unit is used for registering new users. The user login unit is communicated with the user authentication unit for authenticating the user when the user logs in, and determines the target permission information of the user according to the authentication result, and allocates the corresponding system menu to the user according to the target permission information. The permission management unit is communicated with the user authentication unit for allocating the corresponding target permission to the user according to the target permission information after the user completes registration. The log management unit is communicated with the user login unit for recording the user's operation status and generating a record log after the user logs in.
3. The corpus operation platform for industrial large models according to claim 2 is characterized in that: The registration process of the user in the user registration unit includes: After the user enters registration information according to the registration interface, performing activation verification on the user according to the registration information; After the verification is passed, the basic information of the user is collected, and the basic information is parsed for permissions to obtain the corresponding target permission information; The target permission information is sent to the permission management unit, and the corresponding target permission is allocated to the registered user through the permission management unit.
4. The corpus operation platform for industrial large models according to claim 3 is characterized in that: The performing permission parsing on the basic information to obtain corresponding target permission information includes: Converting the basic information into a standard format to obtain standard information, inputting the standard information into a database, and updating a user information table in the database, wherein the user information table is used to store the processed standard information of each user; Calculate the weight coefficient of each data indicator according to the updated user information table; The target weight information of the currently registered user is calculated according to the weight coefficient and the standard information.
5. The corpus operation platform for industrial large models according to claim 4 is characterized in that: The calculating the weight coefficient of each data indicator according to the updated user information table includes: Calculating the mean and standard deviation of the data indicators corresponding to each user in the user information table, wherein the data indicators of the user include at least one of corpus call information, corpus download information, corpus usage frequency and recharge information; Performing a first standardization process on the indicator information according to the mean and the standard deviation to obtain a first standard user information table; Constructing a covariance matrix based on the data in the first standard user information table; Performing eigenvalue decomposition on the covariance matrix to obtain an eigenvalue matrix and an eigenvector matrix; Calculate the load coefficient corresponding to the data indicator of each user according to the eigenvector matrix; Calculate a first weight of each indicator information according to the load coefficient and the eigenvalue matrix; Performing a second standardization process on the data indicators in the user information table to obtain a second standard user information table; Calculate the weight information of each user in the second standard user information table on different data indicators; Obtaining a difference coefficient corresponding to each of the data indicators according to the weight information, and calculating a second weight corresponding to each of the data indicators according to the difference coefficient; A weight coefficient corresponding to the data indicator is calculated according to the first weight and the second weight.
6. The corpus operation platform for industrial large models according to claim 2 is characterized in that: The authority management unit is also used to perform security specification verification on the user at random time during the user's use, and adjust the user's authority according to the security specification verification result; Wherein, when the number of errors in the security specification verification result of the user reaches a first preset number, the authority level of the user is correspondingly reduced; When the number of consecutive correct security specification verification results of the user reaches a second preset number, the authority level of the user is increased.
7. The corpus operation platform for industrial large models according to claim 1 is characterized in that: The corpus management module includes a corpus directory unit, a tool directory unit, and a service provider directory unit. The corpus directory unit is used to generate a corpus directory according to the target authority information of the user. The tool directory unit is communicated with the corpus directory unit and is used to generate a matching tool directory according to the corpus directory. The tool directory is used for the user to perform different processing on the corpus directory. The service provider directory is used to provide service provider information corresponding to each corpus according to the corpus directory.
8. The corpus operation platform for industrial large models according to claim 7 is characterized in that: The process of the tool catalog unit generating a matching tool catalog according to the corpus catalog includes: Acquire first usage information of a tool corresponding to each corpus in the corpus catalog, and calculate first ranking information of a matching tool corresponding to each corpus according to the first usage information, wherein the matching tool is a recommended tool for the corpus; Acquire second usage information of the remaining tools in the existing toolbox, and calculate second ranking information of each of the tools according to the second usage information; The first sorting information corresponding to each of the corpora is added to the second sorting information respectively and then re-sorted to generate target sorting information, and a tool catalog corresponding to each of the corpora is generated according to the result of the target sorting information.
9. The corpus operation platform for industrial large models according to claim 8 is characterized in that: The obtaining of first usage information of each corpus-corresponding tool in the corpus catalog, and calculating first ranking information of each corpus-corresponding matching tool according to the first usage information, includes: Obtaining a first usage frequency, a first evaluation score, a first update frequency, a first accuracy, and a first single usage duration of a matching tool corresponding to each corpus in the corpus directory, and converting the first usage frequency, the first evaluation score, the first update frequency, the first accuracy, and the first single usage duration information into first standard information; Performing weighted calculation on the first standard information according to a first preset weight value to obtain first ranking information of the matching tool corresponding to each of the corpora; The obtaining of the second usage information of the remaining tools in the existing toolbox and calculating the second ranking information of each of the tools according to the second usage information includes: Obtaining the second usage frequency, second evaluation information, second update frequency, second accuracy, and second single usage duration of the remaining tools in the existing toolbox, and converting the second usage frequency, the second evaluation information, the second update frequency, the second accuracy, and the second single usage duration into second standard information; Performing weighted calculation on the second standard information according to the second preset weight value to obtain second ranking information of each of the tools; The step of adding the first sorting information corresponding to each of the corpora to the second sorting information respectively to re-sort and generate target sorting information, and generating a tool catalog corresponding to each of the corpora according to the result of the target sorting information, includes: The first sorting information of the matching tool corresponding to each of the corpora is added to the second sorting information respectively and re-sorted, and the target sorting information corresponding to each corpus is obtained according to the sorting result; A tool catalog corresponding to each of the corpora is generated according to the target sorting information.
10. The corpus operation platform for industrial large models according to claim 7, characterized in that: It is used to provide detailed information of the service provider corresponding to the corpus, sort the detailed information of the service provider according to the sorting result of the corpus, and establish a one-to-one mapping relationship between the corpus and the service provider information. The corpus production module is also communicated with the corpus tracing module to generate the corresponding target corpus according to the provided service provider information.
Citation Information
Patent Citations
Corpus directory management method and system for industrial large model
CN119903124A
Method, device, and computer-readable medium for colormetry considering external environment
KR102596914B1
Secure transmission and exchange of standardized data
US20100064349A1
Generating Models for Text-Dependent Speaker Verification
US20170236520A1
Electric power artificial intelligence model system and working method
WO2025081995A1