A method and device for building a multi-dimensional large language model capability framework
By constructing a multi-dimensional large language model capability framework based on CHC theory and the FLASK system, the limitations of capability definition in existing technologies are solved, high-quality data screening and systematic evaluation of model capabilities are achieved, and the performance of large language models in diverse and specific scenarios is improved.
Patent Information
- Application Number
- CN202510383331.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Existing technologies have limited definitions of the capabilities of large language models, lack multi-dimensional perspectives, and lack systematic exploration of the application of capability frameworks, resulting in low accuracy in data filtering.
Based on the CHC theoretical model, language capabilities are screened, domain-related and non-core capabilities are removed, and cognitive dimension capabilities are refined into pattern recognition, concept abstraction and hypothesis generation. Domain and task dimension capabilities are defined by combining the FLASK domain classification system, and a multi-dimensional large language model capability framework is constructed. High-quality fine-tuning data is obtained by training and labeling the model using GPT-4o and Qwen2.5-7B-Base.
It improves the accuracy of data filtering, enabling the acquisition of high-quality fine-tuning data in diverse general and specific scenarios, thereby enhancing the general intelligence performance of large language models and the accuracy of specific tasks.
Smart Images

Figure CN119918585B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of large language model, in particular to a method and device for building a multi-dimensional large language model capability framework. BACKGROUND
[0002] In recent years, the capabilities of large language models have made significant progress and breakthroughs. In particular, the introduction of reinforcement learning and the application of thought chain reasoning have further enhanced their reasoning capabilities. Currently, large language models are widely used in various industries, such as writing, dialogue, and programming, providing convenience for life. As large language models become more and more complex, it is particularly important to accurately assess their underlying capabilities so that researchers can perceive the current capabilities of large language models and further expand the capability boundaries of large language models. Currently, there are a variety of general benchmark tests widely used to evaluate these capabilities. However, many benchmark tests only focus on certain aspects of model capabilities, such as programming, common sense reasoning, or performance on specific tasks. For example, Massive Multitask Language Understanding (MMLU), a large-scale multitask language understanding benchmark, evaluates the academic knowledge of large language models, but ignores dimensions such as code generation. And the capability dimensions of existing benchmark tests are usually task-oriented, and there is no comprehensive framework to systematically classify and understand the overall capabilities of large language models, making the current industry lack a comprehensive and systematic understanding of the capabilities of large language models.
[0003] In recent years, there have been many studies exploring the capabilities of large language models. Some studies focus on model evaluation, defining evaluation criteria from the aspects of domain, capability, and difficulty, measuring model-generated responses, and evaluating model capabilities. Some studies do not predefine capabilities, but rely on the model itself to label data with capability labels. Some studies approach from the perspective of cognition, dividing model capabilities into four aspects: language knowledge, formal knowledge, world modeling, and social modeling, and constructing test sets from the above four aspects to evaluate model capabilities. Some studies explore the relationship between individual capabilities and cross-capabilities of models. The above studies either focus on measuring model capabilities in a single or limited dimension, such as only defining tasks, or lack detailed definitions and decoupling of capabilities, and the definition of capabilities of large language models is not systematic and comprehensive. In addition, most studies define capabilities for the evaluation of large language models, and lack exploration of the application of capability frameworks. SUMMARY
[0004] To solve the technical problems of the prior art that the definition of the capabilities of large language models is relatively limited, lacks a multi-dimensional perspective, and lacks exploration of the application of capability frameworks, leading to inaccurate data selection, the present application provides a method and device for building a multi-dimensional large language model capability framework. The technical solution is as follows:
[0005] In an aspect, a method for building a multi-dimensional large language model capability framework is provided. The method is implemented by a device for building a multi-dimensional large language model capability framework. The method comprises:
[0006] S1, screening a language capability based on a multi-modal human cognitive capability of a CHC theoretical model; removing a field-related capability and a non-core capability of a large language model based on the language capability, to preliminarily obtain a cognitive dimension capability of the large language model; and refining an induction capability in the preliminarily obtained cognitive dimension capability of the large language model into a pattern recognition capability, a concept abstraction capability, and a hypothesis generation capability, to finally define a cognitive dimension capability of a large prediction model;
[0007] S2, defining a field dimension capability of the large language model based on a FLASK field classification system; and defining a task dimension capability of the large language model;
[0008] S3, constructing a multi-dimensional large language model capability framework according to the cognitive dimension capability of the large language model, the field dimension capability of the large language model, and the task dimension capability of the large language model;
[0009] S4, obtaining a capability annotation training set, performing fine-grained capability label annotation on each instruction in the capability annotation training set by using a GPT-4o model, and obtaining an annotated data set;
[0010] S5, training the multi-dimensional large language model capability framework by using a Qwen2.5-7B-Base model according to the annotated data set, obtaining a trained multi-dimensional capability annotation model, obtaining fine-tuning data of a large language model to be screened, inputting the fine-tuning data of the large language model to be screened into the trained multi-dimensional capability annotation model, and obtaining high-quality fine-tuning data.
[0011] Optionally, the multi-dimensional large language model capability framework is represented by the following formula (1):
[0012] (1)
[0013] wherein, represents the multi-dimensional large language model capability framework; c represents a cognitive capability; d represents a field; and t represents a task. represents a cognitive dimension. represents a field dimension. represents a task dimension.
[0014] Optionally, the formula of the cognitive dimension is represented by the following formula (2):
[0015] (2)
[0016] wherein, represents a specific cognitive ability; represents the total number of cognitive dimensions;
[0017] wherein, the formula of the field dimension is represented by the following formula (3):
[0018] (3)
[0019] wherein, represents a subdivided field; represents the total number of field dimensions;
[0020] wherein, the formula of the task dimension is represented by the following formula (4):
[0021] (4)
[0022] wherein, represents a defined task; represents the total number of task dimensions.
[0023] Optionally, the S4 adopts the GPT-4o model to label each instruction in the ability-labeled training set with a fine-grained ability label, and obtains a labeled data set, including:
[0024] According to the ability-labeled training set, two cognitive dimension ability labels are labeled for each piece of data by the GPT-4o model, one field dimension ability label is labeled for each piece of data by the GPT-4o model, and one task dimension ability label is labeled for each piece of data by the GPT-4o model; by labeling the cognitive dimension ability, field dimension ability and task dimension ability of each piece of data, a labeled data set is obtained.
[0025] Optionally, the trained multi-dimensional ability labeling model includes: a trained cognitive dimension ability labeling model, a trained field dimension ability labeling model, and a trained task dimension ability labeling model.
[0026] Optionally, the fine-tuning data of the large language model to be screened is obtained; the fine-tuning data of the large language model to be screened is input into the trained multi-dimensional ability labeling model to obtain high-quality fine-tuning data, including:
[0027] The fine-tuning data of the large language model to be screened is input into the trained multi-dimensional ability labeling model, and by labeling the cognitive dimension ability, field dimension ability and task dimension ability, a labeled data pool is obtained;
[0028] According to the labeled data pool, all composite abilities in the labeled data pool are defined; wherein, the composite ability is a combination of the cognitive dimension ability, the field dimension ability and the task dimension ability;
[0029] Fine-tuning data screening is performed according to all composite capabilities in the annotation data pool to obtain high-quality fine-tuning data.
[0030] Optionally, the process of defining all composite capabilities in the annotation data pool is represented by the following formula (5):
[0031] (5)
[0032] wherein, represents all composite capabilities in the annotation data pool; represents the annotation data pool; represents the composite capability set of the annotation data pool.
[0033] In another aspect, a device for building a multi-dimensional large language model capability framework is provided. The device is applied to a method for building a multi-dimensional large language model capability framework. The device comprises:
[0034] A first definition unit is configured to filter out language capabilities based on multi-modal human cognitive capabilities of a CHC theoretical model, remove field-related capabilities and non-core capabilities of a large language model based on the language capabilities, and preliminarily obtain cognitive dimension capabilities of the large language model. Based on the preliminarily obtained cognitive dimension capabilities of the large language model, the inductive capabilities are refined into pattern recognition capabilities, concept abstraction capabilities, and hypothesis generation capabilities, and finally the cognitive dimension capabilities of the large language model are defined.
[0035] A second definition unit is configured to define field dimension capabilities of the large language model based on a FLASK field classification system, and define task dimension capabilities of the large language model.
[0036] A construction unit is configured to construct a multi-dimensional large language model capability framework according to the cognitive dimension capabilities of the large language model, the field dimension capabilities of the large language model, and the task dimension capabilities of the large language model.
[0037] An annotation unit is configured to obtain a capability annotation training set, perform fine-grained capability label annotation on each instruction in the capability annotation training set using a GPT-4o model, and obtain an annotated data set.
[0038] A screening unit is configured to train the multi-dimensional large language model capability framework using a Qwen2.5-7B-Base model according to the annotated data set, obtain a trained multi-dimensional capability annotation model, obtain fine-tuning data of a large language model to be screened, input the fine-tuning data of the large language model to be screened into the trained multi-dimensional capability annotation model, and obtain high-quality fine-tuning data.
[0039] Optionally, the multi-dimensional large language model capability framework is represented by the following formula (1):
[0040] (1)
[0041] wherein, represents a multi-dimensional large language model capability framework; c represents cognitive ability; d represents domain; t represents task; represents a cognitive dimension; represents a domain dimension; represents a task dimension.
[0042] Optionally, the formula of the cognitive dimension is represented by the following formula (2):
[0043] (2)
[0044] wherein, represents a specific cognitive ability; represents the total number of cognitive dimensions;
[0045] wherein, the formula of the domain dimension is represented by the following formula (3):
[0046] (3)
[0047] wherein, represents a subdivided domain; represents the total number of domain dimensions;
[0048] wherein, the formula of the task dimension is represented by the following formula (4):
[0049] (4)
[0050] wherein, represents a defined task; represents the total number of task dimensions.
[0051] Optionally, the labeling unit is configured to:
[0052] According to the ability labeled training set, two cognitive dimension ability labels are labeled for each piece of data by the GPT-4o model, one domain dimension ability label is labeled for each piece of data by the GPT-4o model, and one task dimension ability label is labeled for each piece of data by the GPT-4o model; through the cognitive dimension ability, domain dimension ability and task dimension ability labeling of each piece of data, a labeled data set is obtained.
[0053] Optionally, the trained multi-dimensional capability labeling model comprises: a trained cognitive dimension capability labeling model, a trained domain dimension capability labeling model, and a trained task dimension capability labeling model.
[0054] Optionally, the screening unit is configured to:
[0055] input the fine-tuning data of the large language model to be screened into the trained multi-dimensional capability labeling model, and obtain a labeling data pool by labeling the cognitive dimension capability, the field dimension capability and the task dimension capability;
[0056] According to the labeling data pool, all composite capabilities in the labeling data pool are defined; wherein the composite capability is a combination of the cognitive dimension capability, the field dimension capability and the task dimension capability;
[0057] According to all the composite capabilities in the labeling data pool, fine-tuning data screening is performed to obtain high-quality fine-tuning data.
[0058] Optionally, the process of defining all the composite capabilities in the labeling data pool is represented by the following formula (5):
[0059] (5)
[0060] wherein, represents all the composite capabilities in the labeling data pool; represents the labeling data pool; represents the composite capability set of the labeling data pool.
[0061] On the other hand, a device for building a multi-dimensional large language model capability framework is provided, and the device comprises a processor and a memory, wherein the memory stores computer readable instructions, and the computer readable instructions are executed by the processor to implement any one of the methods for building a multi-dimensional large language model capability framework.
[0062] On the other hand, a computer readable storage medium is provided, and the storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement any one of the methods for building a multi-dimensional large language model capability framework.
[0063] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:
[0064] Screening language ability based on the multi-modal human cognitive ability of the CHC theoretical model; removing field-related abilities and non-core abilities of the large language model based on the language ability, to preliminarily obtain the cognitive dimension ability of the large language model; refining the induction ability in the preliminarily obtained cognitive dimension ability of the large language model into the pattern recognition ability, the concept abstraction ability and the hypothesis generation ability, to finally define the cognitive dimension ability of the large language model; defining the field dimension ability of the large language model based on the FLASK field classification system; defining the task dimension ability of the large language model; constructing the multi-dimensional large language model ability framework according to the cognitive dimension ability of the large language model, the field dimension ability of the large language model and the task dimension ability of the large language model; obtaining the ability-labeled training set, adopting the GPT-4o model to label the fine-grained ability tags of each instruction in the ability-labeled training set, to obtain the labeled data set; training the multi-dimensional large language model ability framework by using the Qwen2.5-7B-Base model according to the labeled data set, to obtain the trained multi-dimensional ability labeling model; obtaining the fine-tuning data of the large language model to be screened; inputting the fine-tuning data of the large language model to be screened into the trained multi-dimensional ability labeling model, to obtain high-quality fine-tuning data.
[0065] The application can be applied to general scenario data screening based on diversity to obtain high-quality fine-tuning data in the general scenario, or applied to ability-oriented specific scenario data screening to obtain high-quality fine-tuning data in the specific scenario; the application can improve the accuracy of data screening. BRIEF DESCRIPTION OF DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0067] Figure 1 It is a method flow chart of building a multi-dimensional large language model ability framework provided by the embodiment of the present application;
[0068] Figure 2 It is a whole scheme schematic diagram of the method for building a multi-dimensional large language model ability framework provided by the embodiment of the present application;
[0069] Figure 3 It is an architecture schematic diagram of the cognitive dimension ability, the task dimension ability and the field dimension ability of the method for building a multi-dimensional large language model ability framework provided by the embodiment of the present application;
[0070] Figure 4is a device block diagram of a multi-dimensional large language model capability framework provided by an embodiment of the application.
[0071] Figure 5 is a structural schematic diagram of a device of a multi-dimensional large language model capability framework provided by an embodiment of the application. DETAILED DESCRIPTION
[0072] The technical solutions in the application will be described below with reference to the drawings.
[0073] In the embodiments of the application, the words such as “example”, “for example” and the like are used to represent an example, illustration or description. Any embodiment or design scheme described as “example” in the application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the word “example” is intended to present the concept in a specific manner. In addition, in the embodiments of the application, the meaning expressed by “and / or” can be both, or can be one of the two.
[0074] In the embodiments of the application, “image” and “picture” can be used interchangeably at times. It should be pointed out that the meanings expressed are consistent when the distinction is not emphasized. “Of”, “corresponding” and “corresponding” can be used interchangeably at times. It should be pointed out that the meanings expressed are consistent when the distinction is not emphasized.
[0075] In the embodiments of the application, sometimes the subscript such as W1 can be written in the form of non-subscript such as W1. The meanings expressed are consistent when the distinction is not emphasized.
[0076] In order to make the technical problems, technical solutions and advantages to be solved by the application more clear, the following will be described in detail with reference to the drawings and specific embodiments.
[0077] The embodiments of the application provide a multi-dimensional large language model capability framework building method, which can be implemented by a multi-dimensional large language model capability framework building device. The multi-dimensional large language model capability framework building device can be a terminal or a server. As shown in the multi-dimensional large language model capability framework building method flow chart, the processing flow of the method can include the following steps: Figure 1
[0078] S1, based on the CHC theoretical model of multi-modal human cognitive ability, the language ability is screened out; based on the language ability, the field-related ability and the non-core ability of the large language model are removed, and the cognitive dimension ability of the large language model is preliminarily obtained; based on the preliminarily obtained cognitive dimension ability of the large language model, the induction ability among them is refined into pattern recognition ability, concept abstraction ability and hypothesis generation ability, and finally the cognitive dimension ability of the large language model is defined.
[0079] Among them, based on the CHC theoretical model, a variety of modal human cognitive abilities are covered, including vision, hearing and language expression. The language ability is screened out from the variety of modal human cognitive abilities covered by the CHC theoretical model by manual screening. The memory-related ability and the field-related ability are removed by manual removal. Through screening, the number of CHC-based abilities is reduced from 82 to 14.
[0080] Among them, in order to interface with the large language model, the induction ability is refined, and through refinement, the total number of cognitive dimension abilities of the large language model is changed to 16.
[0081] S2, based on the FLASK field classification system, the field dimension ability of the large language model is defined; the task dimension ability of the large language model is defined.
[0082] Among them, the FLASK field classification system is a field classification system in a fine-grained language model evaluation based on aligned skills.
[0083] Among them, the FLASK field classification system covers 38 specific fields; among them, the business field and the marketing field have similarities, which introduces ambiguity in the ability label model, causing the discrete distribution of labels. Therefore, the original classification system in the application is adjusted from 38 fields to 33 fields by manual means.
[0084] Among them, based on the granularity and integrity of the task, 13 tasks are defined.
[0085] S3, according to the cognitive dimension ability of the large language model, the field dimension ability of the large language model and the task dimension ability of the large language model, a multi-dimensional large language model ability framework is constructed.
[0086] Optionally, the multi-dimensional large language model ability framework is represented by the following formula (1):
[0087] (1)
[0088] Among them, The multi-dimensional large language model ability framework is represented by c, which represents cognitive ability; d represents the field; t represents the task. The cognitive dimension is represented by c. represents the domain dimension; represents the task dimension.
[0089] Optionally, the formula of the cognitive dimension is represented by the following formula (2):
[0090] (2)
[0091] wherein, represents a specific cognitive ability; represents the total number of cognitive dimensions;
[0092] wherein, the formula of the domain dimension is represented by the following formula (3):
[0093] (3)
[0094] wherein, represents a specific domain; represents the total number of domain dimensions;
[0095] wherein, the formula of the task dimension is represented by the following formula (4):
[0096] (4)
[0097] wherein, represents a specific task; represents the total number of task dimensions.
[0098] S4, acquire the ability annotation training set, adopt the GPT-4o model to carry out fine-grained ability label annotation to each instruction in the ability annotation training set, and obtain the annotated data set.
[0099] Optionally, the S4 adopts the GPT-4o model to carry out fine-grained ability label annotation to each instruction in the ability annotation training set, and obtains the annotated data set, including:
[0100] According to the ability annotation training set, two cognitive dimension ability labels are annotated for each data by the GPT-4o model, one domain dimension ability label is annotated for each data by the GPT-4o model, and one task dimension ability label is annotated for each data by the GPT-4o model; through the annotation of the cognitive dimension ability, the domain dimension ability and the task dimension ability of each data, the annotated data set is obtained.
[0101] Wherein, in order to reduce the position bias, when using the GPT-4o model for annotation, the application randomly changes the order of the ability in the prompt word of each data. When annotating cognitive ability, cognitive ability requires in-depth instruction understanding, therefore, the application adopts the GPT-4o model to generate an explanation paired with each label.
[0102] S5, according to the annotated data set, a multi-dimensional large language model capability framework is trained by using a Qwen2.5-7B-Base model, and a trained multi-dimensional capability annotation model is obtained; the fine-tuning data of the large language model to be screened is obtained; the fine-tuning data of the large language model to be screened is input into the trained multi-dimensional capability annotation model, and high-quality fine-tuning data is obtained.
[0103] Optionally, the trained multi-dimensional capability annotation model comprises a trained cognitive dimension capability annotation model, a trained field dimension capability annotation model, and a trained task dimension capability annotation model.
[0104] The cognitive dimension capability annotation model is trained by using the Qwen2.5-7B-Base model to obtain the trained cognitive dimension capability annotation model; the field dimension capability annotation model is trained by using the Qwen2.5-7B-Base model to obtain the trained field dimension capability annotation model; and the task dimension capability annotation model is trained by using the Qwen2.5-7B-Base model to obtain the trained task dimension capability annotation model.
[0105] The FLASK data set is used as the capability annotation training set, the training batch size is 32, the cosine learning rate is set to 2e-5, the full fine-tuning mode is used, a total of 120 steps of fine-tuning are performed, the test set is evaluated once every 40 steps, the annotation of GPT-4o is used as the standard, the accuracy of the model annotation capability is evaluated, and the model corresponding to the highest accuracy is selected as the final model; wherein, 10% of the data in the capability annotation training set is randomly divided as the test set.
[0106] Optionally, the fine-tuning data of the large language model to be screened is obtained; the fine-tuning data of the large language model to be screened is input into the trained multi-dimensional capability annotation model to obtain high-quality fine-tuning data, comprising:
[0107] The fine-tuning data of the large language model to be screened is input into the trained multi-dimensional capability annotation model, and the cognitive dimension capability, the field dimension capability, and the task dimension capability are annotated to obtain an annotation data pool;
[0108] According to the annotation data pool, all composite capabilities in the annotation data pool are defined; wherein, the composite capability is a combination of the cognitive dimension capability, the field dimension capability, and the task dimension capability.
[0109] According to all the composite capabilities in the annotation data pool, the fine-tuning data is screened to obtain high-quality fine-tuning data.
[0110] Optionally, the process of defining all the composite capabilities in the annotation data pool is represented by the following formula (5):
[0111] (5)
[0112] in, Represents all composite capabilities in the annotation data pool; Represents the annotated data pool; A composite capability set representing a labeled data pool.
[0113] In one feasible implementation, a trained multi-dimensional capability annotation model is used to filter data for diverse general scenarios. These scenarios include providing widely applicable capabilities across various tasks, such as conversation, code generation, and content creation, to enhance the general intelligence performance of large language models. The data filtering process for diverse general scenarios includes:
[0114] (1) Obtain a general scenario dataset to be screened; input the general scenario dataset to be screened into the trained multi-dimensional capability annotation model, and obtain a general scenario annotated data pool by annotating cognitive dimension capabilities, domain dimension capabilities, and task dimension capabilities; Based on the general scenario annotated data pool, define all the composite capabilities in the general scenario annotated data pool;
[0115] (2) Select a training dataset from the annotated data pool of general scenarios; based on the selected training dataset, define the composite capability of the training dataset, which is expressed by the following formula (6):
[0116] (6)
[0117] in, Represents the composite capability set of the training dataset; Represents the training dataset.
[0118] (3) Setting a threshold and quantifying diversity by the threshold; wherein the threshold is expressed by the following formula (7):
[0119] (7)
[0120] in, Indicates the cardinality of the set, that is, the number of elements; The value of reflects the coverage of unique composite capabilities in the training dataset relative to the entire annotated data pool;
[0121] (4) By adjusting the set threshold value so that the threshold value is close to 1, the screening index is obtained; according to the screening index, the filtered data is obtained.
[0122] Wherein, when a certain data in the labeled data pool can increase R, the composite ability of the data is added to the composite ability set of the training data set, and the data is added to the training data set; when R cannot be increased, the data corresponding to each composite ability in the labeled data pool is selected to the training data set on average until the training data set reaches the expected data amount.
[0123] In an available implementation, the trained multi-dimensional ability labeling model is used for data screening of specific scenarios based on ability guidance; wherein, the specific scenarios include reading comprehension and medical question answering tasks, so that the performance of the large language model on specific tasks is more accurate and reliable. Wherein, the data screening process of specific scenarios based on ability guidance includes:
[0124] (1) Obtain the validation data set of the specific scenario to be screened; input the validation data set of the specific scenario to be screened into the trained multi-dimensional ability labeling model, and obtain the labeled validation data set by labeling the cognitive dimension ability, the domain dimension ability and the task dimension ability;
[0125] (2) According to the labeled validation data set, a composite ability set of the labeled validation data set is constructed, which is represented by the following formula (8):
[0126] (8)
[0127] Wherein, represents the composite ability set of the labeled validation data set; represents the labeled validation data set.
[0128] (3) Data screening is performed according to the composite ability of the validation data set.
[0129] Wherein, the goal of the data screening process of specific scenarios based on ability guidance is to select the corresponding data of each composite ability in the composite ability set of the labeled validation data set in the labeled data pool on average. In actual operation, the composite ability in the composite ability set of the labeled validation data set may be limited, and the amount of corresponding data of the composite ability in the labeled data pool may not reach the expected data amount. Therefore, the application further decomposes the composite ability in the composite ability set of the labeled validation data set. Specifically, the composite ability in the composite ability set of the labeled validation data set is decomposed into a two-tuple, thereby forming a two-tuple ability set, and the two-tuple ability set is further decomposed into individual dimensions to form a one-tuple ability set; when the composite ability set cannot provide enough data, data screening is performed on the two-tuple ability set, and if the expected data amount is still not reached, data screening is performed on the one-tuple ability set.
[0130] In one possible implementation, the effectiveness of the multi-dimensional large language model capability framework proposed in the present application is verified in the data screening of the general scenario of diversity and the data screening of the specific scenario of capability orientation by using the basic method, the all-data method, the random data method and the INSTAG method.
[0131] The basic method is to directly use the basic large language model for testing without any fine-tuning. The all-data method is to fine-tune the basic model using all the data in the data pool. The random data method is to fine-tune the basic model by randomly selecting data from the data pool. The INSTAG method is an existing large language model capability framework, and the INSTAGGER tagger corresponding to the framework is used to tag the data in the data pool.
[0132] In the data screening of the general scenario of diversity, the general scenario includes: ARC-C contains multiple-choice questions and is aimed at scientific problems for grades 3 to 9; MMLU contains a series of questions about 57 topics, with difficulty levels ranging from elementary to professional, and is mainly used to measure the factual knowledge of large language models; BBH contains a variety of challenging tasks, covering more than 200 subtasks, many of which require higher-order and multi-step reasoning; C-EVAL is a Chinese large language model evaluation framework that covers multiple-choice tasks in multiple fields. As shown in Table 1, the comparison results of the data screening of the general scenario are shown.
[0133] Table 1
[0134]
[0135] As shown in the evaluation results in Table 1, the multi-dimensional large language model capability framework achieved the best overall performance in the four benchmark tests, achieved the best results on MMLU and C-EVAL, and achieved the second best results on BBH and ARC-C. At the same time, even if only a single capability dimension is considered, the method of the present application is superior to the random data method and the INSTAG method, highlighting the accuracy of the multi-dimensional large language model capability framework in defining capabilities and proving the effectiveness of the multi-dimensional large language model capability framework in the data screening of the general scenario of diversity.
[0136] In one possible implementation, in the specific scenario of capability orientation, the large language model needs to have data with specific capabilities to meet the requirements of the corresponding tasks. In the specific scenario, the present application selects three related test sets, each of which represents a specific capability dimension.
[0137] Among them, in the cognitive dimension, the MedQA dataset is selected for testing, which is a medical-related multiple-choice dataset, and experiments are conducted on its subset containing 1,273 test data and 1,272 validation set data. The dataset requires the ability of hypothesis generation and concept abstraction in the cognitive dimension; among them, in the field dimension, four history-related tasks are extracted from MMLU, containing 930 test data and 121 validation set data; among them, in the task dimension, SQuAD is selected as the test task of the task dimension, which is a closed-book question answering and extractive question answering task collected from Wikipedia, containing 10.6k test set data. Since it has no validation set, 200 samples are randomly split from the test set as the validation set. For the validation set with a large amount of data, select up to 200 samples from the validation set of each task for labeling and data screening, and for those validation sets that do not have enough samples, use the entire validation set. As shown in Table 2 is the comparison results of data screening in specific scenarios.
[0138] Table 2
[0139]
[0140] Among them, Table 2 shows the evaluation results of the method of the application under the data screening of specific scenarios. Among them, ACC. represents the accuracy; EM represents the precision matching rate; F1 represents the harmonic mean of precision and recall; the method of the application is better than other methods on the three test sets, and has achieved significant improvement, proving the effectiveness.
[0141] Among them, as Figure 2 is a whole scheme diagram of a multi-dimensional large language model capability framework construction method provided by an embodiment of the application; in a feasible implementation manner, a multi-dimensional large language model capability framework is constructed according to cognitive dimension capabilities, field dimension capabilities and task dimension capabilities; by training the multi-dimensional large language model capability framework, a trained multi-dimensional capability annotation model is obtained; the trained multi-dimensional capability annotation model is applied to multi-diversity-based general scenario data screening to obtain high-quality fine-tuning data in the general scenario; the trained multi-dimensional capability annotation model is applied to capability-oriented specific scenario data screening to obtain high-quality fine-tuning data in the specific scenario.
[0142] Among them, as Figure 3Fig. 1 shows a cognitive dimension capability, a task dimension capability and a field dimension capability architecture of a multi-dimensional large language model capability framework construction method provided by an embodiment of the present application. In a feasible implementation, the cognitive dimension capability defined by the present application includes pattern recognition, concept abstraction, hypothesis generation, general sequence reasoning, quantitative reasoning, communication capability, mathematical performance, reading decoding, writing capability, naming function, associative process, expression process, vocabulary fluency, originality / creativity, idea fluency and problem perception / alternative solution fluency.
[0143] Among them, the task dimension capability defined by the present application includes natural language reasoning, content rewriting, abstract generation, classification, brainstorming, sentiment analysis, completion, generation, bias and fairness, word sense disambiguation, multiple-choice question answering, closed-book question answering and extraction question answering.
[0144] Among them, the field dimension capability defined by the present application includes linguistics, literature, multilingualism, tradition, art, sports, mass media, music, diet, health, biology, earth science, astronomy, chemistry, physics, mathematics, logic, economics, law, politics, pedagogy, sociology, agriculture, computer science, automation, electronics, engineering, programming, communication, religion, philosophy, ethics and history.
[0145] Based on the multi-modal human cognitive ability of the CHC theoretical model, the language ability is screened out; based on the language ability, the field-related ability and the non-core ability of the large language model are removed, and the cognitive dimension capability of the large language model is preliminarily obtained; based on the preliminarily obtained cognitive dimension capability of the large language model, the induction capability is refined into the pattern recognition capability, the concept abstraction capability and the hypothesis generation capability, and finally the cognitive dimension capability of the large language model is defined; based on the FLASK field classification system, the field dimension capability of the large language model is defined; the task dimension capability of the large language model is defined; according to the cognitive dimension capability of the large language model, the field dimension capability of the large language model and the task dimension capability of the large language model, a multi-dimensional large language model capability framework is constructed; an ability annotation training set is obtained, a GPT-4o model is used to perform fine-grained capability label annotation on each instruction in the ability annotation training set, and an annotated data set is obtained; based on the annotated data set, a Qwen2.5-7B-Base model is used to train the multi-dimensional large language model capability framework, and a trained multi-dimensional capability annotation model is obtained; fine-tuning data of a large language model to be screened is obtained; the fine-tuning data of the large language model to be screened is input into the trained multi-dimensional capability annotation model, and high-quality fine-tuning data is obtained.
[0146] The application can be applied to general-scene data screening based on diversity, obtain high-quality fine-tuning data in a general scene, or applied to specific-scene data screening based on ability, obtain high-quality fine-tuning data in a specific scene; and the application can improve the accuracy of data screening.
[0147] Figure 4 is a device block diagram of a multi-dimensional large language model capability framework built according to an exemplary embodiment, and the device is used for a method of building a multi-dimensional large language model capability framework. Referring to Figure 4 , the device includes a first definition unit 410, a second definition unit 420, a construction unit 430, a labeling unit 440, and a screening unit 450. Wherein:
[0148] The first definition unit 410 is configured to filter out language capabilities based on multi-modal human cognitive capabilities of the CHC theoretical model; remove field-related capabilities and non-core capabilities of the large language model based on the language capabilities, and preliminarily obtain cognitive dimension capabilities of the large language model; and based on the preliminarily obtained cognitive dimension capabilities of the large language model, refine the induction capabilities among them into pattern recognition capabilities, concept abstraction capabilities, and hypothesis generation capabilities, and finally define the cognitive dimension capabilities of the large language model.
[0149] The second definition unit 420 is configured to define the field dimension capabilities of the large language model based on the FLASK field classification system; and define the task dimension capabilities of the large language model.
[0150] The construction unit 430 is configured to construct a multi-dimensional large language model capability framework according to the cognitive dimension capabilities of the large language model, the field dimension capabilities of the large language model, and the task dimension capabilities of the large language model.
[0151] The labeling unit 440 is configured to obtain a capability-labeled training set, perform fine-grained capability label labeling on each instruction in the capability-labeled training set by using a GPT-4o model, and obtain a labeled data set.
[0152] The screening unit 450 is configured to train the multi-dimensional large language model capability framework by using a Qwen2.5-7B-Base model according to the labeled data set, obtain a trained multi-dimensional capability labeling model, obtain fine-tuning data of a large language model to be screened, input the fine-tuning data of the large language model to be screened into the trained multi-dimensional capability labeling model, and obtain high-quality fine-tuning data.
[0153] Optionally, the multi-dimensional large language model capability framework is represented by the following formula (1):
[0154] (1)
[0155] wherein, represents a multi-dimensional large language model capability framework; c represents cognitive capability; d represents domain; t represents task; represents a cognitive dimension; represents a domain dimension; represents a task dimension.
[0156] Optionally, the formula of the cognitive dimension is represented by the following formula (2):
[0157] (2)
[0158] wherein, represents a specific cognitive capability; represents the total number of cognitive dimensions;
[0159] wherein, the formula of the domain dimension is represented by the following formula (3):
[0160] (3)
[0161] wherein, represents a subdivided domain; represents the total number of domain dimensions;
[0162] wherein, the formula of the task dimension is represented by the following formula (4):
[0163] (4)
[0164] wherein, represents a defined task; represents the total number of task dimensions.
[0165] Optionally, the labeling unit 440 is configured to:
[0166] According to the capability labeled training set, two cognitive dimension capability labels are labeled for each piece of data by the GPT-4o model, one domain dimension capability label is labeled for each piece of data by the GPT-4o model, and one task dimension capability label is labeled for each piece of data by the GPT-4o model. By labeling the cognitive dimension capability, domain dimension capability and task dimension capability of each piece of data, a labeled data set is obtained.
[0167] Optionally, the trained multi-dimensional capability labeling model comprises a trained cognitive dimension capability labeling model, a trained domain dimension capability labeling model, and a trained task dimension capability labeling model.
[0168] Optionally, the screening unit 450 is configured to:
[0169] The fine-tuning data of the large language model to be screened is input into the trained multi-dimensional capability annotation model, and the cognitive dimension capability, the field dimension capability and the task dimension capability are annotated to obtain an annotation data pool;
[0170] According to the annotation data pool, all composite capabilities in the annotation data pool are defined; wherein the composite capability is a combination of the cognitive dimension capability, the field dimension capability and the task dimension capability;
[0171] According to all the composite capabilities in the annotation data pool, fine-tuning data screening is performed to obtain high-quality fine-tuning data.
[0172] Optionally, the process of defining all the composite capabilities in the annotation data pool is represented by the following formula (5):
[0173] (5)
[0174] wherein, represents all the composite capabilities in the annotation data pool; represents the annotation data pool; represents the composite capability set of the annotation data pool.
[0175] Based on the multi-modal human cognitive capability of the CHC theoretical model, the language capability is screened out; based on the language capability, the field-related capability and the non-core capability of the large language model are removed, and the cognitive dimension capability of the large language model is preliminarily obtained; based on the preliminarily obtained cognitive dimension capability of the large language model, the induction capability therein is refined into the pattern recognition capability, the concept abstraction capability and the hypothesis generation capability, and the cognitive dimension capability of the large language model is finally defined; based on the FLASK field classification system, the field dimension capability of the large language model is defined; the task dimension capability of the large language model is defined; according to the cognitive dimension capability of the large language model, the field dimension capability of the large language model and the task dimension capability of the large language model, a multi-dimensional large language model capability framework is constructed; an ability annotation training set is obtained, a GPT-4o model is used to perform fine-grained capability label annotation on each instruction in the ability annotation training set, and an annotated data set is obtained; according to the annotated data set, a Qwen2.5-7B-Base model is used to train the multi-dimensional large language model capability framework, and a trained multi-dimensional capability annotation model is obtained; fine-tuning data of the large language model to be screened is obtained; the fine-tuning data of the large language model to be screened is input into the trained multi-dimensional capability annotation model, and high-quality fine-tuning data is obtained.
[0176] The application can be applied to general scene data screening based on diversity to obtain high-quality fine-tuning data in the general scene, or applied to capability-oriented specific scene data screening to obtain high-quality fine-tuning data in the specific scene; the application can improve the accuracy of data screening.
[0177] Figure 5 This is a schematic diagram of the structure of a device for building a multi-dimensional large language model capability framework provided by an embodiment of the present invention. Figure 5 As shown, the equipment for building a multi-dimensional large language model capability framework may include the above Figure 4 Optionally, the device 510 for building a multi-dimensional large language model capability framework may include a first processor 2001 .
[0178] Optionally, the device 510 for building a multi-dimensional large language model capability framework may further include a memory 2002 and a transceiver 2003 .
[0179] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.
[0180] The following combination Figure 5 The components of the device 510 that builds the multi-dimensional large language model capability framework are described in detail:
[0181] The first processor 2001 is the control center of the device 510 for building the multi-dimensional large language model capability framework. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).
[0182] Optionally, the first processor 2001 can execute various functions of the device 510 built with the multi-dimensional large language model capability framework by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0183] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 5 CPU0 and CPU1 are shown in FIG.
[0184] In a specific implementation, as an embodiment, the device 510 for building a multi-dimensional large language model capability framework may also include multiple processors, such asFigure 5 1 and 2. The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). A processor herein can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0185] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled by the first processor 2001 for execution. The specific implementation method can refer to the above method embodiment and will not be repeated here.
[0186] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently and be connected to the interface circuit ( Figure 5 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0187] The transceiver 2003 is used to communicate with a network device or a terminal device.
[0188] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 5 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.
[0189] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or can exist independently and be connected to the interface circuit of the device 510 ( Figure 5 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.
[0190] It should be noted that, Figure 5 The structure of the device 510 built by the multi-dimensional large language model capability framework does not constitute a limitation on the router, and the actual knowledge structure recognition device can include more or fewer components than the diagram, or combine certain components, or different component arrangements.
[0191] In addition, the technical effects of the device 510 built by the multi-dimensional large language model capability framework can refer to the technical effects of the method of building a multi-dimensional large language model capability framework described in the above method embodiments, which will not be repeated here.
[0192] It should be understood that the first processor 2001 in the embodiments of the present application can be a central processing unit (CPU), and the processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0193] It should also be understood that the memory in the embodiments of the present application can be volatile or nonvolatile memory, or can include both volatile and nonvolatile memory. The nonvolatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory. The volatile memory can be random access memory (RAM) used as external cache. By way of example, and not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0194] The above-described embodiments can be implemented in whole or in part by software, hardware (such as a circuit), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center through a wired (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0195] It should be understood that the term "and / or" herein is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents that the associated objects before and after it are in an "or" relationship, but it can also represent an "and / or" relationship, which can be understood according to the context before and after it.
[0196] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0197] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined according to their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0198] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0199] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0200] In several embodiments provided by the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0201] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0202] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0203] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0204] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A capability-oriented scenario-specific data filtering method, characterized in that, The specific scenario is a medical question-and-answer scenario, including: Obtaining a verification dataset for a specific scenario to be screened, wherein the verification dataset is a natural language dataset; Input the verification data set of the specific scenario to be screened into the trained multi-dimensional ability annotation model, and obtain the annotated verification data set by annotating the cognitive dimension ability, domain dimension ability, and task dimension ability; According to the annotated validation dataset, a composite capability set of the annotated validation dataset is constructed, which is expressed by the following formula (8): (8) wherein, represents a composite capability set of the annotated validation dataset; represents an annotated validation dataset; data filtering is performed according to the composite capability of the validation dataset; The training process of the multi-dimensional capability labeling model includes: S1. Based on the multimodal human cognitive abilities of the CHC theoretical model, we screen out language abilities. Based on language abilities, we remove domain-related abilities and non-core abilities of the large language model to preliminarily obtain the cognitive dimensional abilities of the large language model. Based on the preliminarily obtained cognitive dimensional abilities of the large language model, we refine the inductive abilities into pattern recognition abilities, conceptual abstraction abilities, and hypothesis generation abilities, and ultimately define the cognitive dimensional abilities of the large prediction model. S2. Based on the FLASK domain classification system, define the domain dimension capabilities of the large language model; define the task dimension capabilities of the large language model; S3. Construct a multi-dimensional large language model capability framework based on the cognitive dimension capability, domain dimension capability, and task dimension capability of the large language model. S4. Obtain the capability labeling training set and use the GPT-4o model to perform fine-grained capability labeling on each instruction in the capability labeling training set to obtain the labeled data set. S5. Based on the labeled data set, the Qwen2.5-7B-Base model is used to train the multi-dimensional large language model capability framework to obtain a trained multi-dimensional capability labeling model, including: a trained cognitive dimension capability labeling model, a trained domain dimension capability labeling model, and a trained task dimension capability labeling model.
2. The method of claim 1, wherein, The multi-dimensional large language model capability framework is expressed by the following formula (1): (1) wherein, represents a multi-dimensional large language model capability framework; c represents cognitive capability; d represents domain; t represents task; represents a cognitive dimension; represents a domain dimension; represents a task dimension.
3. The method of claim 2, wherein, The formula of the cognitive dimension is expressed by the following formula (2): (2) wherein, represents a specific cognitive ability; represents the total number of cognitive dimensions; The formula of the domain dimension is expressed by the following formula (3): (3) wherein, denotes the number of subfields; denotes the total number of field dimensions; The formula of the task dimension is expressed by the following formula (4): (4) wherein, represents a defined task; represents the total number of task dimensions.
4. The method of claim 1, wherein, The S4 uses the GPT-4o model to perform fine-grained capability labeling on each instruction in the capability labeling training set to obtain a labeled dataset, including: According to the ability labeling training set, two cognitive dimension ability labels are labeled for each data through the GPT-4o model, one domain dimension ability label is labeled for each data through the GPT-4o model, and one task dimension ability label is labeled for each data through the GPT-4o model; by labeling the cognitive dimension ability, domain dimension ability and task dimension ability of each data, the labeled data set is obtained.
5. The method of claim 1, wherein, The method further comprises: Obtain fine-tuning data for the large language model to be screened; input the fine-tuning data of the large language model to be screened into the trained multi-dimensional capability annotation model to obtain high-quality fine-tuning data, including: The fine-tuning data of the large language model to be screened is input into the trained multi-dimensional capability labeling model, and the cognitive dimension capability, the field dimension capability and the task dimension capability are labeled to obtain a labeling data pool; According to the labeling data pool, all composite capabilities in the labeling data pool are defined; wherein the composite capability is a combination of the cognitive dimension capability, the field dimension capability and the task dimension capability; According to all the composite capabilities in the labeling data pool, fine-tuning data screening is performed to obtain high-quality fine-tuning data.
6. The method of claim 5, wherein, The process of defining all the composite capabilities in the labeling data pool is represented by the following formula (5): (5) wherein, represents all composite capabilities in the annotation data pool; represents the annotation data pool; represents the composite capability set of the annotation data pool.
7. A capability-oriented scenario-specific data filtering apparatus, characterized by, The specific scene is a medical question and answer scene, and the device is used to implement the data screening method based on the capability orientation of the specific scene according to any one of claims 1-6, characterized in that the device comprises: An acquisition module is configured to acquire a validation data set of a specific scene to be screened; A labeling module is configured to input the validation data set of the specific scene to be screened into the trained multi-dimensional capability labeling model, and label the cognitive dimension capability, the field dimension capability and the task dimension capability to obtain a labeled validation data set; A construction module is configured to construct a composite capability set of the labeled validation data set according to the labeled validation data set, and the construction module is represented by the following formula (8): (8) wherein, represents a composite set of annotated validation data sets; represents an annotated validation data set; A screening module is configured to perform data screening according to the composite capability of the validation data set.
8. A device for data screening based on capability-oriented specific scenarios, characterized in that: The device for data screening of the specific scene based on the capability orientation comprises: A processor; A memory, wherein the memory stores computer readable instructions, and the computer readable instructions are executed by the processor to implement the method according to any one of claims 1 to 6.
9. A computer readable storage medium, characterized in that, The computer readable storage medium stores program code, and the program code can be called and executed by the processor to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Cognitive evaluation method, training method and evaluation system based on psychological measurement network
CN118000684A
Task processing method and device, question and answer processing method and device in target domain, domain task model testing method and device, computing equipment, computer readable storage medium and computer program product
CN118069326A