Data set acquisition method and device and computer equipment
By generating data set requirements and analyzing target parameters, the agent can independently select and obtain the required data sets, solving the problem of limited data sets in the existing technology, improving the diversity and quality of data, and improving the development of AI agents.
Patent Information
- Application Number
- CN202510233189.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The existing technology is difficult to achieve the agent's independent selection and acquisition of required data sets, resulting in limited scale and types of data sets, which limits the development of AI agents.
By generating data set requirements and analyzing target parameters, finding data acquisition experience, planning dataset acquisition methods, or obtaining datasets that meet the needs by conducting data transactions with target agents.
It realizes accurate matching of data set requirements, improves data diversity, quality and compliance, and improves the training effect and performance of large models.
Smart Images

Figure CN120146956A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a method, device, and computer device for obtaining a data set. Background Art
[0002] In recent years, AI (Artificial Intelligence) technology has developed rapidly. As the core tool for implementing AI functions, after the data volume, computing power, and algorithms reach a certain level, the model has gradually evolved into a large model.
[0003] The large model can serve as the core driver of an AI agent (intelligent agent or proxy). With the support of the large model, the AI agent can perform a series of intelligent behaviors such as perception, decision-making, and action. Among them, the basis for the decision-making and action of the AI agent depends on the data set (Data Sets), the training process of the large model depends on the data set, and the data set is an important source for the AI agent to obtain knowledge. Although technicians can manually prepare a certain amount of data sets for the AI agent according to actual needs, the scale and types of manually prepared data sets are often limited. The limited data sets restrict the development of the AI agent. If a method for the AI agent to autonomously obtain data sets can be provided, it will provide strong impetus for the development of the AI agent. Summary of the Invention
[0004] In view of this, the present invention provides a method, device, and computer device for obtaining a data set to solve the problem of how to enable the intelligent agent to autonomously select the required data set.
[0005] In a first aspect, the present invention provides a method for obtaining a data set, which is applied to a large model intelligent agent. The method includes:
[0006] Generating a data set requirement based on the requirements of the current task execution;
[0007] Analyzing the data set requirement to infer a target parameter for characterizing the data set requirement;
[0008] Searching for data acquisition experience based on the target parameter, and using the data acquisition experience to plan a data set acquisition method to obtain the target data set corresponding to the target parameter, and / or obtaining a data set that meets the data set requirement by means of data trading with a target intelligent agent; wherein, the large model intelligent agent is a data purchaser, and the target intelligent agent is a data seller.
[0009] The method for obtaining a dataset provided by an embodiment of the present invention generates a dataset requirement from the requirements of task execution, analyzes the dataset requirement, and infers target parameters, presenting the dataset requirement in the form of quantitative and measurable parameters. This helps to more accurately screen and evaluate data during subsequent searches for data acquisition experience and the process of obtaining a dataset, ensuring that the acquired data highly matches the requirements. Searching for data acquisition experience based on the target parameters and using this experience to plan the dataset acquisition method can save time and resources by drawing on the methods of successfully acquiring similar datasets in the past; and / or obtaining a dataset that meets the requirements through data transactions with target agents, enriching data diversity, improving data quality and comprehensiveness, being beneficial to enhancing the training effect and performance of the large model, and making the data more compliant. The present invention enables the requirements of single-agent and multi-agent scenarios to be considered under a general intelligent agent architecture. One can directly obtain data according to one's own needs or obtain it through transactions, providing a data foundation for the future development of intelligent agents in Continual Learning.
[0010] In an alternative embodiment, generating the dataset requirement based on the requirements of the current task execution includes:
[0011] Generating a dataset requirement according to the requirements of one or more of the current task executions, such as model training tasks, model evaluation tasks, simulation and emulation tasks, knowledge acquisition and update tasks, and tasks of providing data services to users.
[0012] Embodiments of the present invention cover various task types such as model training, evaluation, simulation, knowledge acquisition and update, and data services, enabling the large model intelligent agent to generate corresponding dataset requirements according to the task characteristics regardless of the application scenario, thus facilitating the provision of personalized and customized data services.
[0013] In an alternative embodiment, searching for data acquisition experience based on the target parameters and using the data acquisition experience to plan the acquisition method of the required dataset to obtain the corresponding dataset includes:
[0014] Extracting data acquisition experience matching the target parameters from the memory module, where the memory module is a storage module in the large model intelligent agent for storing historical records of dataset acquisition;
[0015] Determining the dataset acquisition method based on the data acquisition experience matching the target parameters;
[0016] Based on the dataset acquisition method, calling the corresponding data tool in the tool library module and using the data tool to obtain the corresponding dataset.
[0017] In the embodiments of the present invention, data acquisition experiences that match the target parameters are extracted from the memory module. The large model intelligent agent does not need to re-explore the data acquisition method, and can quickly find the past successful experiences for specific target parameters. The intelligent agent can flexibly adjust the acquisition method according to the experiences matched by the parameters to adapt to diverse task requirements, and improve the data acquisition ability and efficiency of the intelligent agent.
[0018] In an alternative embodiment, the method of obtaining a dataset that meets the dataset requirements by means of data trading with a target intelligent agent includes:
[0019] Generating intelligent agent description information based on natural language processing according to the dataset requirements;
[0020] Selecting a target intelligent agent from multiple intelligent agents associated with the multi-intelligent agent data trading platform based on the intelligent agent description information;
[0021] Evaluating the dataset provided by the target intelligent agent and sending a quotation to the target intelligent agent, and based on the preset trading conditions for dataset trading after receiving the feedback of the target intelligent agent's consent to the quotation.
[0022] In the embodiments of the present invention, the target intelligent agent is selected from the multi-intelligent agent data trading platform by using the intelligent agent description information, realizing accurate matching based on requirements, providing rich intelligent agent resources, and enabling data trading between intelligent agents based on clear requirements, reasonable prices and standardized trading conditions, which can stimulate more intelligent agents to participate in data sharing and trading, enrich data resources, and improve the utilization efficiency of data.
[0023] In an alternative embodiment, the method of evaluating the dataset provided by the target intelligent agent and then making a quotation, and based on the preset trading conditions for dataset trading after receiving the feedback of the target intelligent agent's consent to the quotation includes:
[0024] Inputting the dataset provided by the target intelligent agent into a preset valuation model to generate a valuation result, where the preset valuation model is trained by using a preset machine learning model with dataset metadata and trading history records as training data;
[0025] Generating quotation information based on the valuation result and preset trading conditions, and sending it to the target intelligent agent;
[0026] Receiving the feedback information of the target intelligent agent based on the quotation information, and adjusting the quotation information based on the feedback information to generate new quotation information until dataset trading is carried out based on the preset trading conditions after receiving the feedback of the target intelligent agent's consent to the quotation.
[0027] In the embodiments of the present invention, the preset valuation model based on dataset metadata and transaction history records can more comprehensively and accurately evaluate the value of the dataset provided by the target agent, generate relatively scientific and reasonable valuation results, generate quotation information based on the valuation results and preset transaction conditions, so that the quotation not only considers the value of the dataset itself, but also takes into account the transaction needs and restrictions of the large model agent itself, and adjusts according to the feedback information of the target agent based on the quotation information, which improves the flexibility of transaction negotiation and the efficiency of transaction conclusion.
[0028] In an alternative embodiment, the method of obtaining a dataset that meets the dataset requirements by means of data trading with a target agent includes:
[0029] Generating dataset requirement tender information based on dataset requirement information and preset transaction conditions;
[0030] Sending the dataset requirement tender information to a multi-agent data trading platform;
[0031] Obtaining dataset information fed back by at least one agent associated with the multi-agent data trading platform in response to the dataset requirement tender information;
[0032] Screening the fed-back dataset information based on preset evaluation indicators, and trading the datasets that meet the evaluation indicators.
[0033] In the embodiments of the present invention, by publishing tender information on a multi-agent data trading platform, a large number of potential data supply agents can be contacted, the data acquisition channels can be broadened, and the opportunity to find high-quality and suitable datasets can be increased. The tender information attracts multiple agents to respond. When the large model agent screens datasets based on preset evaluation indicators among numerous responses, it can comprehensively consider data quality and transaction costs and select the dataset with the highest cost performance for trading.
[0034] In an alternative embodiment, the trading of datasets that meet the conditions based on preset evaluation indicators includes:
[0035] Extracting the content of each sub-index in the fed-back dataset information, including: information description of the dataset, price of the dataset, delivery method, transaction time, after-sales service, and data qualification certification materials;
[0036] Evaluating each sub-index content respectively using corresponding preset evaluation indicators to obtain the evaluation results of each sub-index;
[0037] Obtaining the comprehensive evaluation results corresponding to each agent based on the evaluation results of each sub-index;
[0038] Based on the comprehensive evaluation results, the shortlisted datasets are obtained and further interacted with the corresponding agents to determine the final target datasets for trading.
[0039] The embodiment of the present invention provides comprehensive, detailed and scientific decision-making support for the large model agent in the data trading process by selecting eligible datasets for trading based on preset evaluation indicators, which helps to improve the quality and efficiency of data trading.
[0040] In an alternative embodiment, the method further includes:
[0041] Generate a trading record;
[0042] Perform a quality assessment on the dataset obtained through trading according to its availability to obtain a quality assessment result.
[0043] The embodiment of the present invention is beneficial for subsequent reference to the trading record by generating the trading record, clearly understanding each link of the transaction, providing detailed basis for solving problems or summarizing experience, and can also optimize the trading strategy through the analysis of a series of trading records. By evaluating the quality of the data obtained through trading, it is ensured that the dataset can meet the actual needs of the large model agent's current task, and the quality assessment result can be fed back into the trading mechanism to help the large model agent adjust the cooperation method or trading conditions with the target agent.
[0044] In an alternative embodiment, the performing a quality assessment on the obtained dataset according to its availability to obtain a quality assessment result includes:
[0045] Perform an availability assessment on the obtained dataset, including: document and description assessment, data format and compatibility assessment, data volume and cost assessment, data quality and compliance assessment;
[0046] When the availability of the obtained dataset meets the dataset requirements and is directly available, a preset evaluation method is used to perform a quality assessment to generate a quality assessment result;
[0047] When the availability of the obtained dataset cannot meet the dataset requirements and is directly available, after processing the obtained dataset based on a preset data annotation tool or a preset enhancement model, a preset evaluation method is used to perform a quality assessment to generate a quality assessment result. When the dataset availability cannot meet the requirements for direct use, it is processed based on a preset data annotation tool or a preset enhancement model, and then the quality is evaluated, which has great flexibility and practicability, so that even data in a poor initial state has the opportunity to meet the usage requirements through processing.
[0048] In the embodiments of the present invention, by carefully evaluating usability in multiple dimensions, it helps to discover various potential problems in the data in advance. When the usability of the obtained dataset meets the requirements and can be directly used, a preset evaluation method is adopted to perform quality evaluation to generate results, which can efficiently complete the determination of data quality. This enables the large model agent to quickly input the data into subsequent tasks, saving time and resources.
[0049] In an alternative embodiment, the method further includes: updating the acquisition process information of the dataset and the quality evaluation result to the memory module, where the acquisition process information of the dataset includes: the process information of obtaining the corresponding dataset using a data tool, and the transaction records of data transactions with the target agent.
[0050] Updating the acquisition process information of the dataset and the quality evaluation result to the memory module in the embodiments of the present invention has the advantages of optimizing the data acquisition strategy, enhancing the decision-making ability of the agent, and promoting the knowledge inheritance and sharing among multiple agents.
[0051] In an alternative embodiment, the dataset requirements include: data volume, data type and format, data quality requirements, delivery time, data usage rights, data security level requirements, and data update frequency expectations.
[0052] By including comprehensive and detailed dataset requirements, it provides multi-dimensional and precise guidance for the data acquisition and use of the large model agent, ensuring that the data fits the task requirements and improving the data utilization efficiency and task execution effect.
[0053] In an alternative embodiment, the target parameters include: data type, data quantity, data distribution, and source, where the data type includes: multi-modal raw data and sensor data, labeled tag data, enhanced or generated data, trained model data, and data used by the agent itself.
[0054] In the embodiments of the present invention, by clearly defining the target parameters, the large model agent can accurately express its own data requirements, thereby more precisely locating and obtaining the required data to meet diverse task requirements.
[0055] In a second aspect, the present invention provides a trading device for a dataset, including:
[0056] A requirements generation module, configured to generate dataset requirements based on the requirements of the current task execution;
[0057] An analysis and reasoning module, configured to analyze the dataset requirements to infer target parameters for characterizing the dataset requirements;
[0058] A data acquisition module, configured to find data acquisition experience based on the target parameter, and to obtain a target data set corresponding to the target parameter by using a data set acquisition method planned by the data acquisition experience, and / or to obtain a data set that meets the data set requirement by way of data transaction with a target agent; wherein, the large model agent is a data purchaser, and the target agent is a data seller.
[0059] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the data set acquisition method according to the first aspect or any corresponding embodiment thereof.
[0060] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the data set acquisition method according to the first aspect or any corresponding embodiment thereof.
[0061] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the data set acquisition method according to the first aspect or any corresponding embodiment thereof. Description of the Drawings
[0062] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0063] Figure 1 is a flowchart of a data set acquisition method according to an embodiment of the present invention;
[0064] Figure 2 is a process diagram of another data set acquisition method according to an embodiment of the present invention;
[0065] Figure 3 is a process diagram of yet another data set acquisition method according to an embodiment of the present invention;
[0066] Figure 4 is a structural block diagram of a data set trading device according to an embodiment of the present invention;
[0067] Figure 5 is a hardware structural diagram of a computer device according to an embodiment of the present invention. Detailed Embodiments
[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0069] In the field of artificial intelligence, an AI Agent (agent or intelligent agent) can be defined as an AI entity that can perceive the environment, make decisions, take actions, and has autonomy, reactivity, pro-activeness, and social ability.
[0070] If a large model is used as the "brain", with the natural language interaction and reasoning capabilities of the large model, a system equipped with an AI Agent can perform corresponding tasks in various forms such as single Agent, multi-Agent, or human-computer interaction Agent, and through capabilities such as autonomous reasoning and planning, tool use, memory retention, and multi-Agent collaboration; taking the large language model (LLM, Large Language Model) as an example, after the large language model matured, the large language model agent (LLM Agent) has seen an explosive development.
[0071] However, there are obvious defects in the current design and framework of AI Agents. For example, the acquisition and trading capabilities of the training data sets required for different models and / or different tasks are not fully considered. Whether an AI Agent provides data services to meet user requirements or based on its own needs for error correction, continuous evolution, upgrade, migration, and generalization, it needs to obtain the required data sets according to the downstream training task requirements.
[0072] Self-acquire or obtain through trading various types of data sets, such as multi-modal raw data sets and sensor data sets, label data sets after human annotation, data sets enhanced, generated, or distilled using traditional or AI algorithms, and model weights, feature data sets, etc. for federated learning.
[0073] In the related art, Information Bazaar is a solution proposed in the field of AI Agent data interaction, where different agents act as information purchasers and sellers respectively. Under this mechanism, the information purchasers who have obtained user authorization play the role of initiating data acquisition requirements. They will post tenders on the Bulletin Board to indicate their needs for specific datasets to the sellers. After receiving the tender information, the sellers will accordingly submit quotes, that is, provide the datasets they own and the terms of the transaction. The information purchasers can repeatedly ask questions to the quoting sellers until they get satisfactory answers or exhaust the authorized budget. After obtaining all the answers, the information purchasers summarize all the answers and feedback them to the user for judgment on whether to purchase. Although Information Bazaar proposes a data trading platform based on multiple agents and realizes the data trading function between different agents to a certain extent, it has obvious limitations. This solution only starts from the perspective of the trading platform and does not fully utilize the capabilities of the agents themselves, restricting the autonomy and flexibility of the agents in the process of data acquisition and trading.
[0074] In order to enable the agent to flexibly and autonomously obtain the required dataset according to actual needs, in this embodiment, a method for obtaining a dataset is provided, which can be applied to large model agents. Figure 1 It is a flowchart of the method for obtaining a dataset according to an embodiment of the present invention. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer device such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here. As Figure 1 shown, the process includes the following steps:
[0075] Step S101, generate a dataset requirement based on the requirements of the current task execution.
[0076] In the embodiment of the present invention, the dataset requirement is generated according to the requirements of the current task execution to ensure that the data is closely related to the task. Specifically, the dataset requirement can be generated according to the requirements of one or more of the current task executions such as model training tasks, model evaluation tasks, simulation and emulation tasks, knowledge acquisition and update tasks, and tasks of providing data services for users. By covering multiple task types, the large model agent can generate corresponding dataset requirements according to the task characteristics regardless of the application scenario, which is conducive to providing personalized and customized data services. The corresponding dataset requirements include: data volume, data type and format, data quality requirements, delivery time, data usage rights, data security level requirements, and expected data update frequency. The following provides some examples for illustration:
[0077] For example, when the current task is a model training task, assume that a binary classification model for predicting customer churn is to be trained and applied to an Internet service platform. The corresponding dataset requirements are as follows:
[0078] 1) Data volume: Considering the sufficiency of model training and the diversity of customer behavior, at least the customer data for the past 3 years needs to be collected, and the expected data volume reaches 100,000 records to ensure that the model can learn enough feature patterns to accurately predict customer churn.
[0079] 2) Data types and formats:
[0080] Basic customer information: including name, age, gender, registration time, etc. The name is of string type, the age is of integer type, the gender is of enumeration type (male / female), and the registration time is in date and time format. It is stored as a CSV file as a whole for convenient data reading and processing.
[0081] Customer behavior data: such as the number of logins, usage duration, purchase frequency, consumption amount, etc. The number of logins and purchase frequency are of integer type, the usage duration is of floating-point type (unit: hour), and the consumption amount is of floating-point type (unit: yuan). The storage format is also CSV.
[0082] Customer churn label: Mark whether the customer has churned, which is a binary classification label (0: not churned, 1: churned), and is integrated with the above data in the same CSV file.
[0083] 3) Data quality requirements:
[0084] Integrity: The data missing rate of each field should be less than 5%. For missing values, they need to be reasonably filled according to business logic. For example, if the age is missing, the average value of the same age group can be used for filling, and if the consumption amount is missing, it can be estimated based on the user level and historical consumption habits.
[0085] Accuracy: The error rate between the basic customer information, behavior data and the actual situation shall not exceed 3% to ensure that the data truly reflects customer behavior.
[0086] Consistency: The statistical caliber of the data needs to be consistent. For example, the statistics of the consumption amount should cover all payment channels and the calculation method should be unified.
[0087] Delivery time: In order not to affect the project progress, the dataset needs to be delivered within 30 days for timely model training work.
[0088] 4) Data usage rights: The data is only for internal use by the Internet service platform for the training and optimization of the customer churn prediction model, and shall not be used for other commercial purposes nor disclosed to third parties.
[0089] 5) Data security level requirements: Since it involves customer personal information and business data, the data security level is set to high. The data needs to be encrypted for storage and transmission, and access to the data requires a strict authentication and authorization process.
[0090] 6) Expected data update frequency: In order to enable the model to reflect changes in customer behavior in a timely manner, it is expected that the data will be updated once a month, and the new data should be collected and organized according to the same data quality requirements.
[0091] For example, when the current task is a model evaluation task, evaluating the above-trained customer churn prediction model, the corresponding dataset requirements are:
[0092] 1) Data volume: Select an independent evaluation dataset with a quantity of 20,000 records. This dataset should be representative and able to reflect the performance of the model in actual applications.
[0093] 2) Data type and format: Consistent with the training data, including customer basic information, customer behavior data, and customer churn labels, all stored as CSV files to ensure data compatibility between the evaluation environment and the training environment.
[0094] 3) Data quality requirements:
[0095] Completeness: The missing rate of each field data should be less than 3%, and the handling method for missing values is consistent with the training data.
[0096] Accuracy: The error rate shall not exceed 2% to ensure high-quality data for accurate model performance evaluation.
[0097] Independence: The evaluation data should be independent of the training data and not contain any records in the training data to ensure the objectivity of the evaluation results.
[0098] Delivery time: Deliver within 5 days after the model training is completed for timely model evaluation.
[0099] 4) Data usage rights: Only used for the evaluation of this customer churn prediction model and shall not be used for other purposes. Similarly, the data confidentiality principle must be strictly observed and not leaked to unauthorized personnel.
[0100] 5) Data security level requirements: The same as the training data security level, set to high. The same encryption storage, transmission measures, and strict access control policies are adopted.
[0101] 6) Expected data update frequency: Since the evaluation data is mainly used for a one-time evaluation of the current model performance, there is no need to update when there are no major changes in the model architecture or training data. If the model is significantly adjusted, the evaluation dataset needs to be prepared again according to the above requirements.
[0102] Step S102: Analyze the dataset requirements to infer the target parameters for characterizing the dataset requirements.
[0103] Specifically, a large number of dataset requirement samples and the corresponding target parameters determined in actual applications can be collected as training data. The samples should cover various different types of tasks and data requirement situations. After training the model, the agent inputs the new dataset requirements into the model, and the model makes inference and prediction to output the corresponding target parameters; or set up agents with different functions, such as domain knowledge agents, data processing agents, model evaluation agents, etc. Information interaction and collaboration are carried out among the agents, and finally a reasonable data volume target parameter is determined. The above process of analyzing and inferring the dataset requirements is only an example and is not limited thereto.
[0104] In an alternative embodiment, the inferred target parameters include: data type, data quantity, data distribution, and source. Among them, the data type includes: multi-modal raw data and sensor data, labeled tag data, enhanced or generated data, trained model data, and data used by the agent itself; the data distribution is, for example, the time distribution range, the target object distribution range, etc. By analyzing the dataset requirements in the embodiments of the present invention to obtain the target parameters, the essential requirements of the task for data can be deeply explored. The following provides some examples of data types for illustration:
[0105] 1. Multi-modal raw data and sensor data, including text, vision, sound, other modalities, and sensors (such as satellite data, robot control data, wireless spectrum data, medical health device data, etc.).
[0106] 2. Labeled tag data, such as data for image classification, object localization, segmentation, etc. in ImageNet, human feedback data in RLHF, video segmentation information in SAM, etc.
[0107] 3. Enhanced or generated data, for example, data enhanced or generated by traditional or AI algorithms, including data generated by using traditional methods for data cleaning, filling in blanks, complementing, etc., or data generated by large models.
[0108] 4. Trained model data, such as model weight data that has been trained locally, feature data after training, etc.
[0109] 5. Data used by the agent itself, such as communication records, data in the agent's memory module, tool data, etc., including text, vision, sound, other modalities, and sensor data, etc.
[0110] Step S103: Based on the target parameters, search for data acquisition experience, and use the data acquisition method planned by the data acquisition experience to obtain the target data set corresponding to the target parameters, and / or obtain the data set that meets the data set requirements by means of data trading with the target agent; wherein, the large model agent is the data purchaser, and the target agent is the data seller.
[0111] In the embodiment of the present invention, by extracting the matching historical experience from the memory module, the process of re-exploring the data acquisition path is avoided, saving a large amount of time and effort. Based on these experiences, the data source and acquisition method can be selected to ensure that the acquired data has high quality and meets the task requirements. At the same time, the combination of multiple data acquisition methods, such as the combination of public data sets and web scraping, can increase the diversity of data and make the model training more comprehensive and accurate. The experience in the memory module can help select the most economical and effective data acquisition method, giving priority to using free public data sets. For the data that needs to be supplemented additionally, apply for open source data sets (such as MIMIC, etc.) through the web form filling tool or email tool in the tool library module. This acquisition method can save time and economic costs.
[0112] In addition to obtaining data through its own planning, the embodiment of the present invention can also obtain the corresponding data set through transactions with multiple agents associated with the multi-agent data trading platform, which can provide diverse data sources for the large model agent. The target agent may have unique data sets that are difficult to obtain by other means. Through data trading, the data required for large model training can be supplemented, enriching the data diversity, helping to improve the generalization ability and robustness of the model, and the data sets obtained through trading are more compliant with data regulations.
[0113] In practical applications, according to the actual needs of the agent, various factors such as cost, data diversity, and data compliance can be comprehensively considered to select one of the two data set acquisition methods, or through the cooperation or complementary application of the data obtained by the two acquisition methods.
[0114] In one embodiment, the process of searching for data acquisition experience based on the target parameters in step S103 and using the data acquisition method planned by the data acquisition experience to obtain the target data set corresponding to the target parameters specifically includes the following steps:
[0115] A1. Extract the data acquisition experience that matches the target parameters from the memory module, where the memory module is a storage module in the large model agent for storing the historical records of data set acquisition.
[0116] Specifically, taking the example of training an image classification model to identify different breeds of cats: Assume the target parameters are that the image format is JPEG, the resolution is not less than 500x500 pixels, and it contains at least 10 different breeds of cats, with at least 1000 images for each breed.
[0117] The memory module stores the historical records of previously obtained image datasets. By retrieving the memory module, it is found that when training a pet image recognition model before, there was experience in obtaining a similar cat image dataset, including: searching for relevant cat image datasets on public image dataset platforms (such as ImageNet), which have a large amount of labeled image data and high data quality. Using web crawler tools to scrape cat pictures uploaded by users from pet forums and social media platforms (such as pet-related topics on Instagram).
[0118] A2. Based on the data acquisition experience that matches the target parameters, determine the dataset acquisition method.
[0119] Specifically, according to the above acquisition experience, determine the following dataset acquisition methods: Give priority to downloading data from public dataset platforms because this data has been sorted and labeled to a certain extent, and the acquisition efficiency is high. At the same time, use web crawler tools to scrape additional data from pet forums and social media platforms to supplement images of cats in different scenarios and angles, increasing the diversity of the data.
[0120] A3. Based on the dataset acquisition method, call the corresponding data tools in the tool library module and use the data tools to obtain the corresponding dataset.
[0121] Specifically, for downloading data from public dataset platforms, call the download tools in the tool library. For example, for scraping data from the network, call the web crawler tools in the tool library to obtain the required dataset.
[0122] In one embodiment, when obtaining a dataset that meets the dataset requirements by means of data trading with the target agent in step S103, the dataset providing agent can be found in an active or passive manner for data trading.
[0123] In one embodiment, the active method includes the following steps:
[0124] B1. Based on the dataset requirements, generate agent description information in a natural language processing manner.
[0125] In one example, the dataset requirements cover behavioral data such as browsing records, purchase records, search keywords, and dwell time of at least 1 million active users on the platform in the past two years. The data format is CSV, the data missing rate does not exceed 5%, and the data update frequency is once a month. Through natural language processing technology, the above requirements are transformed into agent description information: "Find an agent that has multi-dimensional behavioral data of more than 1 million active users on the e-commerce platform in the past two years. The data includes browsing, purchase, search keyword, and dwell time records, the format is CSV, the missing rate is less than 5%, and the data can be updated monthly."
[0126] B2, based on the agent description information, select a target agent from multiple agents associated with the multi-agent data trading platform;
[0127] For example, after receiving this description information, the multi-agent data trading platform filters among the many agents it is associated with. For instance, there are agents in the platform that focus on financial data and educational data, which do not meet the requirements and are excluded. Eventually, an agent that has been collecting and organizing e-commerce user behavioral data for a long time is selected. This agent claims to have a large-scale user behavioral data that meets the conditions and can ensure data quality and update frequency, and it is determined as the target agent.
[0128] B3, conduct an evaluation of the dataset provided by the target agent and feedback a quotation to the target agent. After receiving the feedback from the target agent agreeing to the quotation, conduct a dataset transaction based on the preset transaction conditions;
[0129] Specifically, the e-commerce enterprise agent evaluates the data sample provided by the target agent, considering factors such as data scale, quality, update frequency, and its own budget, and gives a quotation. Suppose after evaluation, the e-commerce enterprise believes that this batch of data is worth 500,000 yuan and feedbacks this quotation to the target agent. After receiving the quotation, the target agent conducts an internal evaluation, deems the quotation reasonable, agrees to the quotation and feedbacks it to the e-commerce enterprise. At this time, both parties conduct a transaction according to the preset transaction conditions. The preset transaction conditions may include data delivery method (such as transmission through an encrypted network), delivery time (within one week), data usage rights (only for building a user purchase behavior prediction model, not for resale or other purposes), etc. After meeting these conditions, the target agent delivers the dataset to the e-commerce enterprise to complete this data transaction.
[0130] In one embodiment, the above step B3 specifically includes the following steps:
[0131] B31, input the dataset provided by the target agent into a preset valuation model to generate a valuation result. The preset valuation model is trained using the dataset metadata and transaction history records as training data and a preset machine learning model;
[0132] For example, after the target agent provides a data sample, an e-commerce enterprise inputs it into a preset valuation model. The training data of this model includes the metadata of numerous past datasets (such as data scale, data dimension, data quality evaluation indicators, etc.) and relevant transaction history records (transaction price, situation of both parties in the transaction, etc.), and is trained using a preset machine learning model such as a random forest. The model analyzes the metadata of the dataset provided by the target agent. For example, the data scale reaches the behavior data of 1.5 million active users, the dimension is rich and contains various behavior records, and the data quality is detected with a missing rate of only 3%. Combining with the transaction prices of similar data transactions in the past, the final generated valuation result is 600,000 yuan.
[0133] B32. Generate a quotation message based on the valuation result and preset transaction conditions, and send it to the target agent;
[0134] For example, according to the valuation result of 600,000 yuan, combined with the preset transaction conditions (such as the data needs to be delivered within 5 days, 50% of the payment is made first after delivery, and the remaining payment is made after one month of trouble-free data use, etc.), a quotation message is generated. The quotation is 550,000 yuan, clearly informing the target agent of transaction conditions such as the delivery time and payment method, and then sending it to the target agent.
[0135] B33. Receive the feedback message from the target agent based on the quotation message, and generate a new quotation message after adjusting the quotation message based on the feedback message, until after receiving the feedback of the target agent's consent to the quotation, conduct the dataset transaction based on the preset transaction conditions.
[0136] For example, after receiving the quotation, the target agent feedbacks that the price is too low and the expected price is 620,000 yuan. The purchasing agent receives the feedback message, re-evaluates its own urgency for data needs, budget ceiling, and the prices of similar data in the market, and adjusts the quotation message. The adjusted new quotation is 580,000 yuan and is sent to the target agent again. The target agent still hopes the price can reach 600,000 yuan after the second feedback. After weighing, the purchasing agent agrees to raise the price to 600,000 yuan. At this time, receiving the feedback of the target agent's consent to the quotation, both parties follow the preset transaction conditions. The target agent delivers the data through an encrypted network within 5 days, and the purchasing agent pays 300,000 yuan first and pays the remaining 300,000 yuan after one month of trouble-free use, completing the dataset transaction.
[0137] The valuation model trained based on the dataset metadata and transaction history records in the embodiments of the present invention can scientifically and reasonably value the dataset by integrating various factors, generate a quotation message by combining the valuation result and preset transaction conditions, make the quotation include comprehensive transaction elements, and continuously adjust the quotation according to the feedback of the target agent, reflecting the flexibility and negotiation of the transaction process.
[0138] In one embodiment, the passive method includes the following steps:
[0139] C1. Generate dataset requirement tender information based on dataset requirement information and preset transaction conditions;
[0140] For example, generate the dataset requirement tender information as "Now inviting bids to purchase user behavior datasets, which are required to include browsing, purchase, search keywords, and stay time records of over 1 million active users in the past two years, in CSV format, with a missing rate ≤ 5%, and updated monthly. The transaction conditions are delivery within 7 days, 40% of the payment is made first after delivery, and the balance is paid after one month of use without problems. The data is only used for building prediction models and cannot be resold. Eligible agents are welcome to participate in the bidding."
[0141] C2. Send the dataset requirement tender information to a multi-agent data trading platform.
[0142] Specifically, send the dataset requirement tender information to a multi-agent data trading platform, which has many associated agents involved in data collection and trading in different fields.
[0143] C3. Obtain dataset information feedback by at least one agent associated with the multi-agent data trading platform in response to the dataset requirement tender information.
[0144] Specifically, for example, after receiving the tender information, the multi-agent data trading platform pushes it to the associated agents and obtains dataset information feedback by multiple agents. Agent A indicates that it has behavior data of 1.2 million active users in the past three years, with a data missing rate of 2%, which can meet monthly updates, but the delivery time requires 10 days; Agent B states that it has behavior data of 1.5 million active users in the past two years, with a missing rate of 3%, and can be delivered within 7 days, but has more stringent restrictions on data usage rights. In addition to being used for prediction models, its supervision is required during the model training process.
[0145] C4. Screen the feedback dataset information based on preset evaluation indicators and conduct transactions on the datasets that meet the evaluation indicators.
[0146] In the embodiment of the present invention, by publishing tender information on a multi-agent data trading platform, a large number of potential data supply agents can be contacted, the data acquisition channels can be broadened, and the opportunity to find high-quality datasets can be increased. When the large model agent screens datasets based on preset evaluation indicators among numerous responses, it can comprehensively consider data quality and transaction costs and select the dataset with the highest cost performance for transactions.
[0147] In one embodiment, the above step C4 specifically includes:
[0148] C41. Extract the content of each sub-index in the feedback dataset information, including: information description of the dataset, price of the dataset, delivery method, transaction time, after-sales service, and data qualification certification materials;
[0149] Specifically, the information description of the dataset covers key information such as data volume, time range, data quality, and update frequency, enabling the purchaser to intuitively understand the scale and quality level of the data. For example, by knowing that Agent A has the behavior data of 1.2 million active users in the past three years, with a missing rate of 2% and updated monthly, it can be judged whether the data is sufficient to support model training and whether it meets the requirements of time span and data quality, ensuring the applicability of the data.
[0150] By clarifying the price of the dataset, it is possible to evaluate whether the cost is reasonable in combination with one's own budget and market conditions. When comparing the data provided by different agents, price is an important consideration factor, and the dataset with the highest cost performance is selected on the premise of meeting the requirements.
[0151] By understanding the delivery method, such as whether it is encrypted transmission and whether breakpoint resumption is supported, the security and integrity of the data during transmission can be ensured, preventing data leakage and damage. At the same time, based on the transaction time, an agent that can deliver data in a timely manner can be selected according to the project progress requirements, ensuring the smooth progress of the project.
[0152] By clarifying the content of after-sales service, such as response time and problem-solving ability, if there are quality problems or doubts about the data, the agent can be required to respond and solve quickly according to the after-sales service terms, reducing the impact of data problems on the business.
[0153] By checking the data qualification certification materials, such as whether there is data certification by an authoritative agency, the legality of the data source and the reliability of the data can be verified, helping to avoid the risk of using illegal or low-quality data and ensuring the accuracy and compliance of the data for model training and business analysis.
[0154] C42. Evaluate each sub-index content respectively using the corresponding preset evaluation index to obtain the evaluation results of each sub-index;
[0155] In the embodiments of the present invention, through preset evaluation indexes, such as evaluating the data volume, data time range, data quality (missing rate), data update frequency, etc. in the information description of the dataset, the value of the data for its own needs can be accurately judged. For example, for the data of Agent A and B, after evaluation, it can be known that the data volume of Agent B is slightly larger, while the time range of Agent A is wider, which can help clarify which data is more suitable for its own model training and business analysis needs.
[0156] C43. Obtain the corresponding comprehensive evaluation results of each agent based on the evaluation results of each sub-index.
[0157] Specifically, assume that the price, delivery method, after-sales service, and data qualification certification materials make Agent A and Agent B perform equally well. Considering only the currently known information, Agent B has advantages in terms of data volume and transaction time, while Agent A has an advantage in terms of data time range. After comprehensive consideration, the comprehensive evaluation result of Agent B is slightly higher than that of Agent A.
[0158] C44. Based on the comprehensive evaluation results, obtain the shortlisted data sets, and conduct further interactions with the corresponding agents to determine the final target data sets for transactions.
[0159] Specifically, for example, if it is determined that the data set of Agent B is shortlisted, further interact with Agent B on details such as price, delivery method, after-sales service, and data qualification certification materials. After multiple rounds of interactions, it is determined that the final price is 800,000 yuan, the delivery method is encrypted network transmission, the after-sales service includes free data problem consultation and solution within one year, and the data qualification certification materials are complete and certified by an authoritative institution. After both parties reach an agreement, it is determined that the data set of Agent B is the final target data set for transactions.
[0160] Based on the above steps S101 - step S103, the method for obtaining a data set provided by an embodiment of the present invention, as Figure 2 shown, further includes:
[0161] Step S104, generate a transaction record; by generating a transaction record in an embodiment of the present invention, it is beneficial for subsequent review of the transaction record, clearly understanding each link of the transaction, providing detailed basis for problem-solving or experience summary, and also optimizing the transaction strategy through the analysis of a series of transaction records.
[0162] Step S105, conduct a quality assessment on the data set obtained through the transaction according to its availability to obtain a quality assessment result. In one embodiment, step S105 specifically includes the following steps:
[0163] D1. Conduct an availability assessment on the obtained data set, including: document and description assessment, data format and compatibility assessment, data volume and cost assessment, data quality and compliance assessment;
[0164] For example, check whether the data set provided by the agent is accompanied by detailed documents, including data source description, data collection method, data field meaning explanation, etc., and evaluate whether it meets the business requirements for documents and descriptions; confirm whether the data format is the format required by the tender, and whether this format is compatible with existing data processing tools and analysis software. For data volume and cost assessment, for example, evaluate whether the data volume meets the requirements for building a user purchase behavior prediction model, and at the same time judge whether the cost is reasonable in combination with the purchase price; check whether the data missing rate is the value promised by the agent, whether there are outliers in the data, and whether the data complies with relevant laws, regulations and privacy policies.
[0165] D2. When the availability of the obtained dataset meets the direct availability requirement of the dataset, a preset evaluation method is adopted to perform quality evaluation to generate a quality evaluation result.
[0166] For example, the preset evaluation method includes evaluating aspects such as the accuracy, integrity, consistency, and timeliness of the data. In terms of accuracy, by comparing with some known accurate data, it is found that the data accuracy rate reaches 98%; in terms of integrity, each field and record are checked, and no key data is found missing; in terms of consistency, the values of fields with the same meaning in the data are consistent in different records; in terms of timeliness, the data update frequency is once a month, which can meet the requirements of data timeliness. Based on the comprehensive evaluation of various items, a quality evaluation result is generated.
[0167] D3. When the availability of the obtained dataset cannot meet the direct availability requirement of the dataset, after processing the obtained dataset based on a preset data annotation tool or a preset enhancement model, a preset evaluation method is adopted to perform quality evaluation to generate a quality evaluation result.
[0168] Specifically, for example, the availability of the obtained dataset does not meet the requirements. For example, the data format is the XLS format that is not commonly used by the enterprise, and the data missing rate is as high as 10%. At this time, the obtained dataset is processed based on a preset data annotation tool or a preset enhancement model. The data format conversion tool is used to convert the XLS format to the CSV format, and the data filling algorithm and the data enhancement model are used to fill and enhance the missing data, so that the data missing rate is reduced to less than 5%. After processing, the preset evaluation method is adopted for quality evaluation, and the evaluation result shows that the data quality reaches the available standard and meets the basic requirements for building the model.
[0169] By carefully evaluating the availability in multiple dimensions in the embodiments of the present invention, it helps to discover various possible problems of the data in advance. When the availability of the obtained dataset meets the requirements and can be directly used, a preset evaluation method is adopted to perform quality evaluation to generate a result, which can efficiently complete the determination of data quality. This enables the large model agent to quickly input the data into subsequent tasks, saving time and resources.
[0170] Based on the above steps S101 - step S105, the method for obtaining a dataset provided by the embodiments of the present invention, as Figure 3 shown, further includes:
[0171] Step S106, updating the dataset acquisition process information and the quality evaluation result to the memory module, where the dataset acquisition process information includes: the process information of obtaining the corresponding dataset using a data tool and the transaction record of data transactions with the target agent.
[0172] Specifically, the embodiments of the present invention save the process information of obtaining a data set using data tools, such as which download tools, crawler tools, and specific parameter settings are used. When encountering a similar data acquisition requirement next time, it is possible to quickly refer to past successful experiences, skip the cumbersome tool selection and parameter debugging processes, and directly reuse effective methods, greatly improving the data acquisition efficiency. When it is necessary to obtain user behavior data again, a data acquisition process can be quickly built based on the tool usage experience recorded in the memory module.
[0173] Record the data transaction records with the target agent, including transaction prices, delivery times, after-sales service terms, etc., which helps to evaluate the reputation and reliability of different agents in future data transactions. By analyzing past transaction records, high-quality agents that deliver on time and have stable data quality can be identified, avoiding cooperation with agents with poor reputations and reducing transaction risks. At the same time, by comparing transaction prices in different periods, the market conditions can be better grasped, and a more reasonable price can be obtained in subsequent transactions, effectively controlling the data procurement cost.
[0174] Recording the quality assessment results in the memory module is conducive to comparative analysis of the data quality of different batches. If the subsequent model training effect is not good, it is possible to trace back to the corresponding data acquisition process and quality assessment results to find the root cause of the data quality problem, such as whether it is caused by incomplete data cleaning, excessive data augmentation, etc., which helps to optimize the data processing process targeted, improve the data quality, and thus improve the model performance and the accuracy of business decisions.
[0175] In addition, in the multi-agent system or different version iterations of the agent, the data acquisition process information and quality assessment results in the memory module can be passed on as important knowledge. Different agents can share relevant information in the memory module to achieve knowledge exchange and collaboration. For example, if an agent has rich experience in obtaining data sets in a specific field, by sharing this information, other agents can learn from its methods and jointly improve the level of data acquisition and management, promoting the development of the entire multi-agent system.
[0176] In this embodiment, a data set acquisition device is also provided. This system is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0177] This embodiment provides a data set acquisition device, which is applied to as Figure 4 shown, and includes:
[0178] A requirements generation module 401 for generating dataset requirements based on the requirements of the current task execution.
[0179] An analysis and reasoning module 402 for analyzing the dataset requirements to infer target parameters for characterizing the dataset requirements.
[0180] A data acquisition module 403 for finding data acquisition experiences based on the target parameters, and for using the data acquisition experiences to plan the dataset acquisition method, to acquire the target dataset corresponding to the target parameters, and / or to acquire the dataset that meets the dataset requirements by means of data trading with the target agent; where the large model agent is the data purchaser and the target agent is the data seller.
[0181] In some alternative embodiments, the requirements generation module 401 includes generating dataset requirements according to the requirements of one or more of the current task executions among model training tasks, model evaluation tasks, simulation and emulation tasks, knowledge acquisition and update tasks, and tasks of providing data services for users.
[0182] In some alternative embodiments, the data acquisition module 403 includes:
[0183] A data experience extraction unit for extracting data acquisition experiences matching the target parameters from the memory module, where the memory module is a storage module in the large model agent for storing historical records of dataset acquisition.
[0184] A data acquisition method generation unit for determining the dataset acquisition method based on the data acquisition experiences matching the target parameters.
[0185] A first dataset acquisition unit for invoking the corresponding data tool in the tool library module based on the dataset acquisition method and using the data tool to acquire the corresponding dataset.
[0186] In some alternative embodiments, the data acquisition module 403 includes:
[0187] An agent description information generation unit for generating agent description information based on the dataset requirements in a natural language processing manner.
[0188] A target agent screening unit for selecting the target agent from multiple agents associated with the multi-agent data trading platform based on the agent description information.
[0189] A first trading unit for valuing the dataset provided by the target agent and providing a feedback quote to the target agent, and for conducting dataset trading based on the preset trading conditions after receiving the feedback of the target agent's agreement to the quote.
[0190] In some alternative embodiments, the first trading unit includes:
[0191] An evaluation subunit, configured to input the dataset provided by the target agent into a preset valuation model to generate a valuation result. The preset valuation model is trained by using a preset machine learning model with dataset metadata and trading history records as training data;
[0192] An offer subunit, configured to generate offer information based on the valuation result and preset trading conditions, and send it to the target agent;
[0193] A first trading subunit, configured to receive feedback information from the target agent based on the offer information, adjust the offer information based on the feedback information to generate new offer information, and perform dataset trading based on the preset trading conditions until receiving the consent offer feedback from the target agent.
[0194] In some alternative embodiments, the data acquisition module 403 includes:
[0195] A tender information generation unit, configured to generate dataset requirement tender information based on dataset requirement information and preset trading conditions;
[0196] A tender information sending unit, configured to send the dataset requirement tender information to the multi-agent data trading platform;
[0197] A response information acquisition unit, configured to acquire dataset information fed back by at least one agent associated with the multi-agent data trading platform in response to the dataset requirement tender information;
[0198] A second trading unit, configured to screen the fed-back dataset information based on preset evaluation indicators, and trade the datasets that meet the evaluation indicators.
[0199] In some alternative embodiments, the second trading unit includes:
[0200] An index extraction subunit, configured to extract the content of each sub-index in the fed-back dataset information, including: information description of the dataset, price of the dataset, delivery method, trading time, after-sales service, and data qualification certification materials;
[0201] An evaluation subunit, configured to evaluate each sub-index content by using corresponding preset evaluation indicators respectively to obtain the evaluation results of each sub-index;
[0202] A comprehensive evaluation result generation subunit, configured to obtain the corresponding comprehensive evaluation results of each agent based on the evaluation results of each sub-index;
[0203] A second trading subunit, configured to obtain a shortlisted dataset based on the comprehensive evaluation result, further interact with the corresponding agent, and determine the final target dataset for trading.
[0204] In some alternative embodiments, the above-mentioned apparatus further includes:
[0205] A trading record module, configured to generate trading records;
[0206] A trading data quality evaluation module, configured to evaluate the quality of the dataset obtained through trading according to its availability, and obtain a quality evaluation result.
[0207] In some alternative embodiments, the trading data quality evaluation module includes:
[0208] An availability evaluation subunit, configured to evaluate the availability of the obtained dataset, including: document and instruction evaluation, data format and compatibility evaluation, data volume and cost evaluation, data quality and compliance evaluation;
[0209] A first quality evaluation result obtaining subunit, configured to, when the availability of the obtained dataset meets the dataset requirements and is directly available, perform quality evaluation using a preset evaluation method to generate a quality evaluation result;
[0210] A second quality evaluation result obtaining subunit, when the availability of the obtained dataset does not meet the dataset requirements and is directly available, after processing the obtained dataset based on a preset data annotation tool or a preset enhancement model, perform quality evaluation using a preset evaluation method to generate a quality evaluation result.
[0211] In some alternative embodiments, the above-mentioned system further includes: a memory module update module, configured to update the dataset acquisition process information and the quality evaluation result to the memory module, where the dataset acquisition process information includes: the process information of obtaining the corresponding dataset using a data tool, and the trading record of data trading with the target agent.
[0212] In some alternative embodiments, the dataset requirements include: data volume, data type and format, data quality requirements, delivery time, data usage rights, data security level requirements, data update frequency expectations.
[0213] In some alternative embodiments, the target parameters include: data type, data quantity, data distribution, source, where the data type includes: multi-modal raw data and sensor data, labeled tag data, enhanced or generated data, trained model data, data used by the agent itself.
[0214] The further function descriptions of the above modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here. The device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0215] The embodiment of the present invention also provides a computer device having the above Figure 4 data set acquisition device shown. Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 5 shown, the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 5 In
[0216] FIG. 10, one processor 10 is taken as an example.
[0217] Among them, the memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments. The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device and the like. In addition, the memory 20 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely disposed relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0218] The memory 20 may include volatile memory, for example, random access memory; the memory may also include non-volatile memory, for example, flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memory. The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0219] An embodiment of the present invention further provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processed on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0220] A part of the present invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0221] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for acquiring a data set, applied to a large model agent, characterized in that: include: Generate data set requirements based on the requirements of the current task execution; Analyzing the data set requirements to infer target parameters for characterizing the data set requirements; Based on the target parameters, data acquisition experience is searched, and a data set acquisition method planned using the data acquisition experience is used to acquire a target data set corresponding to the target parameters, and / or a data set that meets the data set requirements is acquired by conducting data transactions with a target intelligent agent; wherein the large model intelligent agent is a data purchaser and the target intelligent agent is a data seller.
2. The method according to claim 1, characterized in that The data set requirements are generated based on the requirements of the current task execution, including: Generate data set requirements based on the requirements of one or more current tasks including model training tasks, model evaluation tasks, simulation and emulation tasks, knowledge acquisition and update tasks, and tasks of providing data services to users.
3. The method according to claim 1, characterized in that The searching for data acquisition experience based on the target parameter and using the data acquisition experience to plan a method for acquiring the required data set to acquire the corresponding data set includes: Extracting data acquisition experience matching the target parameter from a memory module, wherein the memory module is a storage module in the large model agent for storing historical records of data set acquisition; Determining a data set acquisition method based on the data acquisition experience matching the target parameter; Based on the data set acquisition method, the corresponding data tool in the tool library module is called, and the data tool is used to acquire the corresponding data set.
4. The method according to claim 1, characterized in that The method of acquiring a data set that meets the data set requirements by performing data transaction with the target intelligent agent includes: According to the requirements of the data set, generate agent description information based on natural language processing; Based on the agent description information, a target agent is selected from multiple agents associated with the multi-agent data trading platform; The data set provided by the target intelligent entity is evaluated and a quotation is fed back to the target intelligent entity. After receiving feedback from the target intelligent entity agreeing with the quotation, the data set transaction is carried out based on preset transaction conditions.
5. The method according to claim 4, characterized in that The process of making a quotation after evaluating the data set provided by the target intelligent entity, and trading the data set based on preset trading conditions after receiving the target intelligent entity's agreed quotation feedback, includes: Input the data set provided by the target agent into a preset valuation model to generate a valuation result, wherein the preset valuation model is trained using a preset machine learning model based on the data set metadata and transaction history records as training data; Generate quotation information based on the valuation result and preset transaction conditions, and send it to the target intelligent agent; Receive feedback from the target intelligent agent based on the quotation information, and generate new quotation information after adjusting the quotation information based on the feedback information, until the target intelligent agent agrees to the quotation feedback, and then conduct data set transactions based on preset transaction conditions.
6. The method according to claim 1, characterized in that The method of acquiring a data set that meets the data set requirements by performing data transaction with the target intelligent agent includes: Generate data set demand bidding information based on data set demand information and preset transaction conditions; Sending the data set demand bidding information to the multi-agent data trading platform; Acquire data set information fed back by at least one agent associated with the multi-agent data trading platform in response to the data set demand bidding information; The feedback data set information is screened based on the preset evaluation indicators, and the data sets that meet the evaluation indicators are traded.
7. The method according to claim 6, characterized in that The step of selecting a qualified data set for trading based on a preset evaluation index includes: Extracting the sub-item indicators in the feedback data set information, including: data set information description, data set price, delivery method, transaction time, after-sales service and data qualification certification materials; The corresponding preset evaluation indicators are used to evaluate the content of each sub-item indicator to obtain the evaluation results of each sub-item indicator; Based on the evaluation results of each sub-indicator, the comprehensive evaluation results corresponding to each intelligent agent are obtained; Based on the comprehensive evaluation results, the shortlisted data sets are obtained, and further interactions are carried out with the corresponding intelligent agents to determine the final target data set for trading.
8. The method according to claim 2, characterized in that: The method further comprises: Generate transaction records; The quality of the data set obtained through the transaction is assessed according to its availability to obtain a quality assessment result.
9. The method according to claim 8, characterized in that The quality assessment of the acquired data set is performed according to its availability to obtain a quality assessment result, including: Conduct usability assessment on the acquired data sets, including: documentation and description assessment, data format and compatibility assessment, data volume and cost assessment, data quality and compliance assessment; When the availability of the acquired data set meets the data set requirements and is directly available, a preset evaluation method is used to perform quality evaluation and generate a quality evaluation result; When the availability of the acquired data set cannot meet the data set demand for direct availability, after the acquired data set is processed based on a preset data annotation tool or a preset enhancement model, a quality assessment is performed using a preset assessment method to generate a quality assessment result.
10. The method according to claim 8, characterized in that The method also includes: updating the acquisition process information and quality assessment results of the data set into the memory module, wherein the acquisition process information of the data set includes: process information of obtaining the corresponding data set using data tools, and transaction records of data transactions with the target intelligent entity.
11. The method according to claim 1, characterized in that: The data set requirements include: data volume, data type and format, data quality requirements, delivery time, data usage rights, data security level requirements, and data update frequency expectations.
12. The method according to claim 1, characterized in that The target parameters include: data type, data quantity, data distribution, and source, wherein the data types include: multimodal raw data and sensor data, annotated label data, enhanced or generated data, trained model data, and data used by the agent itself.
13. A data set transaction device, applied to a large model intelligent agent, characterized in that: include: The requirement generation module is used to generate data set requirements based on the requirements of the current task execution; An analysis and reasoning module, used for analyzing the data set requirements to infer target parameters for characterizing the data set requirements; A data acquisition module is used to search for data acquisition experience based on the target parameters, and to use the data acquisition experience to plan a data set acquisition method to acquire a target data set corresponding to the target parameters, and / or to acquire a data set that meets the data set requirements by conducting data transactions with a target intelligent agent; wherein the large model intelligent agent is a data purchaser and the target intelligent agent is a data seller.
14. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method for acquiring a data set according to any one of claims 1 to 12 by executing the computer instructions.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method for obtaining a data set according to any one of claims 1 to 12.
16. A computer program product, characterized in that The method comprises computer instructions, wherein the computer instructions are used to cause a computer to execute the method for obtaining a data set according to any one of claims 1 to 12.
Citation Information
Patent Citations
Learning training platform of automatic valuation model
CN117235524A
Purchase decision processing method and device based on multi-agent collaboration
CN118428750A
Automatic data analysis method and system based on artificial intelligence
CN118626828A
Intelligent agent optimization method, device and system based on reinforcement learning, and storage medium
CN119443200A
Instantiating machine-learning models at on-demand cloud-based systems with user-defined datasets
US20220383150A1