Data value evaluation method and system based on knowledge mining large model and analogue simulation agent
By constructing external knowledge search tools, designing a data value assessment framework, conducting self-reflective reinforcement learning, and simulating intelligent agent game theory, the scientific and consensus-building issues of data value assessment are resolved, assessment efficiency and result reliability are improved, and market-based allocation of data elements is supported.
Patent Information
- Application Number
- CN202610050683.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-02-13
AI Technical Summary
Existing technologies make it difficult to scientifically and rationally assess the value of data, resulting in low efficiency in the market-based allocation of data elements and difficulty in reaching a consensus on value between supply and demand.
This approach employs a knowledge mining-based large-scale model and a simulation-based intelligent agent. It involves constructing an external knowledge search tool, designing a data value assessment framework, conducting self-reflective reinforcement learning training, and using the simulation-based intelligent agent to simulate the game between supply and demand, thereby generating a structured data value assessment report.
It achieves scenario adaptability and result reliability in data value assessment, improves assessment efficiency, supports large-scale market-based allocation of data elements, and ensures consensus and transaction credibility between supply and demand sides.
Smart Images

Figure CN121526720A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, specifically to a data value assessment method and system based on a knowledge mining large model and a simulation intelligent agent, used to automatically and structurally assess the potential value of data in specific application scenarios. Background Technology
[0002] With data being recognized as a new factor of production following land, labor, capital, and technology, the marketization and value realization of data elements has become a key focus of current development. In this process, the scientific assessment of data value is a crucial challenge in promoting the market-based allocation of data elements. This dilemma stems from the inherent uniqueness of data value and the cognitive misalignment between data supply and demand. On the one hand, data value assessment is a typical interdisciplinary and complex task. The same data can have vastly different value potentials depending on different business objectives, technological conditions, and market environments. This characteristic requires the assessment work to deeply integrate data science, domain knowledge, and operational management—a multi-dimensional knowledge system—placing high demands on the comprehensive capabilities of the assessment body. On the other hand, it is difficult for data supply and demand to establish an effective consensus on value. Suppliers often cannot clearly articulate the value path of data in specific scenarios, while demanders, limited by information asymmetry and data usage capabilities, find it difficult to accurately predict the application effects and return on investment of data.
[0003] At the technical methodology level, existing data valuation methods primarily follow the three traditional intangible asset valuation methods: the cost approach, the income approach, and the market approach. However, these methods all exhibit significant shortcomings when applied to data elements: the cost approach fails to reflect the marginal benefits generated by data reuse across various scenarios, and cost collection and measurement are difficult; the income approach is difficult to accurately predict due to the high uncertainty of data value; and the market approach lacks effective comparable benchmarks because the data element market is currently in a "cold start" phase and has not yet formed a stable market foundation. Therefore, existing methods generally have poor applicability in practice and are unable to support large-scale, high-efficiency market-based allocation of data elements.
[0004] Large language models demonstrate powerful natural language understanding and logical reasoning capabilities, providing a new technical approach to solving the challenge of data value assessment. However, directly applying general-purpose large models to the field of data value assessment has significant limitations: First, there is a lack of a systematic value analysis framework. Although general-purpose large models learn from massive amounts of knowledge corpora during training, they struggle to spontaneously and structurally extract the knowledge they have learned. Second, the consistency and verifiability of assessment results are poor. The generative nature of large models leads to randomness in their output; multiple requests for the same assessment scenario may produce different value paths and parameter suggestions. Third, there is a lack of psychological game-playing between data supply and demand. The intrinsic value of data is highly uncertain, context-dependent, and subjective, and cannot be directly anchored. The formation of its final transaction price is essentially the result of a dynamic game between supply and demand parties under conditions of information asymmetry, based on their respective strategies, beliefs, and psychological expectations.
[0005] In conclusion, there is an urgent need in this field for a value assessment method that can overcome the aforementioned shortcomings. This method should be able to integrate professional domain knowledge, ensure the stability and reliability of the output results, and ultimately form a scientific and reasonable value assessment result that can be recognized by both the supply and demand sides. Summary of the Invention
[0006] The purpose of this invention is to provide a data value assessment method and system based on a knowledge mining large model and a simulation intelligent agent. It aims to utilize the knowledge integration and reasoning capabilities of the large model, and through knowledge mining technology and simulation intelligent agents, automatically generate data schemes and assess their value, thereby overcoming the limitations of existing methods and improving the efficiency and accuracy of data value assessment.
[0007] To achieve the above objectives, the present invention provides the following two technical solutions. Firstly, the present invention provides a data value assessment method based on a large knowledge mining model and a simulated intelligent agent. The method includes the following steps:
[0008] 1. Build external knowledge search tools to adapt to different data usage scenarios and provide models with domain knowledge and information needed for value assessment and quantification in those scenarios;
[0009] 2. Design a data value assessment framework, which consists of pre-defined general value dimensions to serve as a systematic framework for guiding value discovery and ensuring the comprehensiveness of value exploration;
[0010] 3. Self-reflection reinforcement learning training. Enhance the model's self-reflection ability, automatically correct common-sense errors in the generation process, and make its answers more logical and consistent with common sense;
[0011] 4. Construct a simulation agent that includes data suppliers, demanders, and a third party for value assessment. Introduce Rubinstein's game theory to simulate the dynamic game process of bargaining under conditions of information asymmetry. The supply and demand agents can engage in complex interactions such as multiple rounds of bidding, counter-bidding, and conditional agreements. At the same time, the third-party agent can simulate the game process under different market positions and market environments by adjusting the states of the supply and demand agents, thus reproducing the psychological changes of the supply and demand parties during negotiations.
[0012] 5. Generate a value quantification assessment report, summarize the list of consensus-based data usage solutions, and automatically output a structured data value assessment report.
[0013] Secondly, the present invention provides a data value assessment system based on a knowledge mining large model and a simulated intelligent agent for implementing the method described in the first aspect above. The system includes five modules corresponding one-to-one with the steps of the method:
[0014] 1. Knowledge mining module, used to search for domain knowledge and information needed for value assessment and quantification, to mine and stimulate the knowledge contained within the large model;
[0015] 2. The value assessment framework module is used to configure and manage general value dimensions preset by experts, providing the system with a structured value assessment perspective;
[0016] 3. Simulation of intelligent agents module, used to simulate the data usage schemes and data value consensus formation process of data supply and demand sides under different market conditions;
[0017] 4. The report generation module is used to summarize the list of consensus-based data usage schemes and generate a structured data value assessment report.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0019] First, the assessment framework is better suited to data value assessment scenarios. The scenario-adaptation method shifts from directly defining value to defining data usage plans, providing a new path for revealing the potential value of data and reaching a consensus.
[0020] Second, the value assessment results are more realistic and reliable. By enhancing generation through retrieval and training through reinforcement learning, the model's illusion is reduced, and its professionalism is improved. Combined with simulation of intelligent agents to simulate real pricing scenarios and their paths.
[0021] Third, automation and high efficiency. It achieves full-process automation from scenario analysis, knowledge retrieval, value reasoning to report generation, improving evaluation efficiency and supporting large-scale market-based allocation of data elements. Attached Figure Description
[0022] Figure 1The flowchart illustrates a data value assessment method based on a knowledge mining large model and a simulated intelligent agent, as provided in this embodiment of the invention.
[0023] Figure 2 This is a framework diagram of a data value assessment system based on a knowledge mining large model and a simulation intelligent agent, provided for an embodiment of the present invention.
[0024] Figure 3 This is a schematic diagram of the composition of a computer system provided in an embodiment of the present invention. Detailed Implementation
[0025] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Please refer to... Figure 1 The specific implementation process of this method is as follows:
[0026] Step S1: Construct an external knowledge search tool. This tool enables the model to overcome the static and limited nature of its internal knowledge. Upon receiving a specific business scenario, it can autonomously and purposefully retrieve and acquire highly relevant domain knowledge, typical data application cases, and market parameters and industry benchmark information required for value assessment and quantification from authoritative external knowledge sources such as industry databases, policy websites, and research institution publishing platforms. This provides timely, accurate, and structured external knowledge input for subsequent value framework analysis, data application scheme generation, and simulation.
[0027] As one implementation method, step S1 can be specifically implemented as the following steps S11-S14:
[0028] Step S11: Search Source Configuration and Strategy Definition. The system manages a scalable search source configuration library, which defines the types of authoritative knowledge sources to be prioritized and their access strategies for different data usage scenarios. This includes, but is not limited to, regulatory agency websites for obtaining industry policies, regulations, standards, and official statistics; industry research institutions and consulting firm platforms for obtaining market analysis reports, industry white papers, and best practice cases; academic databases and standards organizations for obtaining value assessment theories, methodological literature, and technical standards; and leading enterprise technology blogs and case studies for obtaining cutting-edge technology application scenarios and highly effective data usage solutions.
[0029] Step S12: Scenario-driven query construction. When a specific business scenario description is received, the system calls the feature extraction module to perform in-depth analysis of the scenario. External knowledge search tools will utilize these semantic tags to automatically construct accurate and structured search query statements.
[0030] Step S13: Multi-source Asynchronous Retrieval and Content Acquisition. Based on the constructed query and matching knowledge sources, the tool initiates parallel, asynchronous network retrieval requests. It employs a configurable request scheduler, integrating various technologies such as BeautifulSoup for static page parsing, Selenium for dynamic page rendering, and standardized API calls to adapt to the page structures of different knowledge sources. All retrieval actions adhere to the robots.txt protocol, with reasonable request intervals and retry mechanisms, a proxy IP pool for rotation, and anti-anti-crawler strategies to ensure the legitimacy, stability, and friendliness to the target server during the retrieval process.
[0031] Step S14: Query Result Parsing, Deduplication, and Structured Summarization. Obtain the original query results, including HTML pages, JSON data, and PDF documents. Then, use the corresponding parser to extract plain text content from the original format. Utilize named entity recognition and key sentence extraction techniques to quickly extract core information closely related to the query intent from the text. Deduplicate similar information from different sources and sort them comprehensively based on the authority of the information source, the timeliness of the content, and the semantic relevance to the query tags. Organize the filtered core information into one or more structured text summaries, clearly annotating the information source, publication time, and metadata. The summary format facilitates reading and understanding by the large model, serving as key context for the "Query Enhancement Generation" step.
[0032] Step S2: Design a Data Value Assessment Framework. This framework is a thought system designed to systematically guide the data value assessment process. Its core consists of general value dimensions pre-defined by domain experts based on economics, management, and industry practice. These dimensions constitute a complete value perspective coverage system, specifically including the following aspects: First, risk-avoidance value, focusing on the role of data in identifying, warning against, and mitigating potential risks, with the core analysis direction being the potential losses avoided; second, efficiency-enhancing value, focusing on the role of data in optimizing processes and saving resources, with the core analysis direction being the time or cost saved; third, revenue-generating value, focusing on the role of data in expanding markets, increasing sales, and innovating business models, with the core analysis direction being the new revenue directly or indirectly generated; fourth, decision-making optimization value, focusing on the role of data in supporting accurate judgment and scientific planning, with the core analysis direction being the improvement of decision-making quality or speed; and fifth, security and compliance value, focusing on the role of data in enhancing system resilience and meeting regulatory requirements, with the core analysis direction being the reduction of security risks or compliance costs. This template is not used as a quantitative scoring standard, but rather as a structured thinking guide checklist to ensure the breadth of value exploration.
[0033] Step S3: Self-reflective reinforcement learning training. Reinforcement learning is a machine learning method whose core idea is that an agent learns how to take actions to maximize a certain cumulative reward through trial and error by interacting with the environment. Through reinforcement learning training, the large language model used to build the agent can automatically correct common-sense errors in the generation process, making its answers more logical and more in line with common sense.
[0034] Group Relative Policy Optimization (GRPO) is a reinforcement learning algorithm for large language models. GRPO's core advantage lies in its outcome-oriented optimization strategy. For the same input, the model generates N candidate answers and directly calculates the relative advantage between each answer within the group, without relying on an additional value network. This design allows it to effectively optimize its strategy even under sparse reward signals. Value assessment tasks have stringent requirements for the logical rationality and common-sense compliance of the results. Large language models are based on autoregressive generation mechanisms, and their unidirectional prediction characteristics result in a lack of backtracking and verification capabilities for the output content during the generation process, leading to illogical and unrealistic evaluation results. Therefore, this patent attempts to introduce a self-reflection mechanism through the GRPO algorithm to block logical fallacies. However, in extremely sparse reward environments, GRPO's convergence speed is often slow and unstable. This patent improves upon this by using a forced alignment group relative policy optimization algorithm (FA-GRPO) for reinforcement learning training, inserting a forced alignment "reflection trigger" during the Rollout phase. Specifically, after the model initially outputs y1 for the same question, it doesn't immediately reward the user. Instead, it forcibly appends a prompt at the end: "To ensure correctness, I should verify the previous logic again." This prompts the model to continue generating y2. The complete sequence [y1 + trigger + y2] is then treated as a single complete response, and the overall advantage is calculated. Forced alignment and reflection can lead to overfitting. The model might over-reflect on a correct answer, "correcting" it to an incorrect conclusion, or get bogged down in meaningless repetitive reflection, reducing efficiency. To mitigate these risks, this patent designs a dynamic reward mechanism. New penalty and reward terms are added to the reward function. Penalties are imposed for "reflecting a correct answer incorrectly as an error" and "excessive reflection rounds," while rewards are given for "successfully correcting an initial incorrect answer to correct through reflection." Through this training loop, the model is effectively guided to gradually learn prudent reasoning habits, ultimately enabling it to generate logically rigorous, reasonable, and highly credible value assessment conclusions for complex business scenarios.
[0035] Step S4: Formation of Multi-Role Value Consensus Based on Simulated Intelligent Agents. This step aims to construct a simulation environment composed of multiple intelligent agents. By simulating the interaction between data supply and demand parties and third-party evaluators, and combining bargaining game theory with joint value simulation, a consensus-based data usage plan and value range are formed. This process replaces the traditional debate mechanism, simulating the probing, compromise, and agreement reaching in real business negotiations, making the evaluation results more realistic and the transactions more credible.
[0036] As one implementation method, step S4 can be specifically implemented as the following steps S41-S44:
[0037] Step S41: Multi-role Intelligent Agent Initialization and Role Modeling. An intelligent agent is an intelligent entity capable of perceiving the environment, making autonomous decisions, and taking actions to achieve specific goals. The system initializes three independent simulated intelligent agents, representing the data supplier, data demander, and third-party value assessor, respectively. Each agent is built based on a domain-adapted large language model and injected with differentiated role goals, knowledge background, and behavioral strategies through carefully designed system prompts. The supplier agent aims to maximize the perceived value and transaction benefits of the data while protecting its own sensitive data information and business logic. It is endowed with a deep understanding of the data source, quality, and potential application scenarios, but its negotiation strategy may lean towards optimism. The demander agent aims to obtain data application solutions that can effectively solve its business pain points and bring measurable returns while controlling costs and risks. It is endowed with a clear understanding of its own business scenarios, resource constraints, and data usage capabilities, and its negotiation strategy leans towards conservatism and pragmatism. The third-party agent aims to facilitate a fair and reasonable value consensus, acting as a relatively neutral coordinator and evaluator. It is endowed with a deep understanding of value assessment frameworks, industry benchmarks, and market rules, and is responsible for guiding the negotiation process, proposing compromise solutions, and independently calculating the value of the solutions.
[0038] Step S42: Joint Simulation Environment Setup. The designed data value assessment framework, along with relevant domain knowledge, market parameters, and case information acquired through external search tools, is synchronized as a common knowledge background to all agents. The real raw data represented by each agent, its core business logic, and cost structure-sensitive information are not exposed in the simulation environment. Agents reason and make decisions solely based on their role objectives and the common knowledge. Based on this common knowledge, all parties begin to conceive and preliminarily evaluate potential data utilization solutions and their value around specific business scenarios.
[0039] Step S43: Multi-round interactive simulation based on Rubinstein game theory. Rubinstein game theory emphasizes that participants engage in multiple rounds of bidding and counter-bidding based on external choice value and strategic expectations, ultimately reaching equilibrium through gradual compromise. The demand-side agent proposes an upper limit of value and its rationale, the supply-side agent proposes a lower limit of value and its basis, and a third-party agent provides a benchmark assessment based on value dimensions and market parameters. Subsequently, the system drives the agents to conduct multiple rounds of iterative game. In each round, each party can dynamically adjust the details of the solution and its value expectations based on the other party's bid, third-party suggestions, and its own strategy. Their behavioral logic is embedded with Rubinstein game elements. The third-party agent continuously introduces industry benchmarks and quantitative parameters during negotiations, guiding both parties to seek compromise on key value dimensions. The system calculates the consensus level in real time and terminates when the consensus level reaches a preset threshold or the maximum number of simulation rounds. If a consensus is reached, the system outputs a list of data usage solutions and value ranges accepted by both parties; if a complete consensus is not reached, the system outputs the latest solution status and key points of disagreement. All negotiation trajectories, psychological changes, and supporting evidence generated during the simulation are structurally recorded, forming a traceable "negotiation-evaluation" evidence chain. This chain is then directly transmitted as a consensus input to the subsequent value quantification and report generation stages, ensuring that the simulation results are deeply embedded in the overall evaluation process and enhancing the realistic rationality and transaction credibility of the final value conclusion.
[0040] The system records the entire parallel generation process, multiple rounds of debate, and changes in consensus throughout, forming an auditable decision-making chain. This provides support for the interpretability of the final solution and fully demonstrates the value of feature extraction and knowledge mining in solution optimization.
[0041] Step S5: Value Assessment Report Generation. Value assessment report generation is the process of transforming the consensus-based data utilization plan list into quantifiable benefit estimates and automatically generating a final structured assessment report. First, the system automatically analyzes the consensus-based data utilization plan list, negotiation trajectory, and psychological changes. For each path's "expected output / value" stage, it reverse-engineers the key market parameters required to convert it into specific economic value, forming a structured list of value quantification market information. Specifically, the system automatically identifies and outputs the set of parameters required for quantification by analyzing the inherent logic of the value path. This process ensures that each value path has a corresponding, executable quantification plan. Second, the system automatically generates the data utilization plan and value assessment report according to a standardized template. This report contains two core parts: one is the consensus-based data utilization plan list, which details the complete description of each plan, including the plan name, triggering conditions, data input, implementation actions, and structured information on expected benefits; the other is the value quantification parameter list, which clearly presents the set of quantification parameters corresponding to each plan and clarifies the role of each parameter in value calculation. For plans that did not reach a complete consensus during the debate stage, the report retains the latest version and marks the main points of contention, ensuring the completeness of the assessment results. The entire report is output in Markdown structured format, which supports subsequent systematic processing and facilitates review and decision-making by managers.
[0042] Corresponding to the data value assessment method based on knowledge mining large models and simulated intelligent agents described in the first aspect above, this invention also provides a systematic implementation method. Please refer to... Figure 2 , Figure 2 This is a system framework diagram provided for an embodiment of the present invention. The system includes the following modules: a knowledge mining module, used to search for domain knowledge and information required for value assessment quantification, and to mine and stimulate the knowledge contained within the large model; a value assessment framework module, used to configure and manage general value dimensions preset by experts, providing the system with a structured value assessment perspective; a simulation agent module, used to simulate the data usage schemes of data supply and demand parties under different market conditions and the data value consensus formation process; and a report generation module, used to summarize the list of consensus-based data usage schemes and generate a structured data value assessment report.
[0043] Please see Figure 3 , Figure 3This is a schematic diagram of a computer system provided in an embodiment of the present invention. The computer system includes at least a processor 101, a communication interface 102, and a memory 103. The processor 101, communication interface 102, and memory 103 can be connected via a bus or other means. The processor 101 (or central processing unit, CPU) is the computing and control core of the computer system, capable of parsing various instructions and processing various data within the computer system. The communication interface 102 may optionally include a standard wired interface or a wireless interface (such as WIFI, mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for data transmission and interaction within the computer system. The memory 103 is a memory device in the computer system used to store programs and data. It is understood that the memory 103 here can include the computer system's built-in memory, or it can include extended memory supported by the computer system. The memory 103 provides storage space, which stores the computer system's operating system.
[0044] In one embodiment, the processor 101 executes the data value assessment method based on a knowledge mining large model and a simulated intelligent agent provided in the above embodiments of the present invention by running a computer program in the memory 103.
Claims
1. A data value assessment method based on a large knowledge mining model and a simulated intelligent agent, characterized in that, Includes the following steps: Build an external knowledge search tool to retrieve and obtain relevant domain knowledge, typical data usage cases, and market parameters required for value quantification from authoritative external knowledge sources after receiving a specific data usage scenario. Then, parse, deduplicate, and organize the search results into structured text summaries with metadata annotations. Design a data value assessment framework, which consists of pre-defined general value dimensions. This framework serves as a systematic way of thinking to guide value discovery. The general value dimensions include risk aversion value, efficiency improvement value, revenue growth value, decision optimization value, and security and compliance value. The basic large language model is trained by self-reflection reinforcement learning. The forced alignment group relative strategy optimization algorithm FA-GRPO is adopted. After the model generates the initial answer, forced concatenation reflection trigger prompts are used to guide the model to perform logical verification of its own output. A dynamic reward mechanism is used to give positive rewards to reflection behavior that successfully corrects errors and to punish behavior that over-reflects or incorrectly corrects the correct answer, thereby improving the logical consistency and common sense conformity of the generated content. The system constructs a simulation agent, which includes three types: data supplier agent, data demander agent, and value assessment third-party agent. Based on Rubinstein game theory, it simulates a multi-round bargaining process under conditions of information asymmetry. Each agent makes offers, counter-offers, and conditional agreements based on its role objectives, common knowledge background, and negotiation strategies. The third-party agent introduces industry benchmarks and quantitative parameters to guide the negotiation. The system calculates the consensus degree and terminates the simulation when it reaches a preset threshold or the maximum number of rounds, outputting a list of consensus-based data usage schemes and their corresponding value ranges. The consensus-based data usage plan list, negotiation trajectory, and psychological changes are automatically analyzed to deduce the key quantitative parameters required for each plan to realize its economic value. A structured data value assessment report containing a description of the data usage plan and a list of value quantification parameters is automatically generated according to a standardized template.
2. The method according to claim 1, characterized in that, The steps for constructing an external knowledge search tool specifically include: Configure an extensible search source library and define the regulatory websites, industry research platforms, academic databases, and enterprise technology case libraries that are prioritized for querying under different data usage scenarios; Automatically construct structured query statements based on semantic features of input data usage scenarios; Initiate multi-source asynchronous retrieval requests, and integrate static page parsing, dynamic rendering, and API technologies to obtain the original content; The system extracts content, identifies named entities, and extracts key sentences from search results in HTML, PDF, and JSON formats. It then performs deduplication and sorting based on source authority, timeliness, and semantic relevance to generate structured summaries with metadata annotations.
3. The method according to claim 1, characterized in that, The steps for constructing the simulated intelligent agent include: Three independent agents are initialized, representing the data supplier, the data demander, and the third party for value assessment, respectively. Each agent is built based on a large language model trained by self-reflective reinforcement learning, and its differentiated role goals and behavioral strategies are configured through system prompts. The domain knowledge and market parameters obtained by the data value assessment framework and external knowledge search tools are synchronized to all intelligent agents as a public knowledge background. In a simulation environment, the intelligent agents are driven to interact in multiple rounds based on Rubinstein game theory. In each round, each party dynamically adjusts the details of the plan and the expected value based on the other party's bid, the third party's suggestions and its own strategy. The third-party intelligent agent continuously introduces industry benchmarks and value quantification parameters to guide both parties to seek compromise on key value dimensions. When the consensus reaches a preset threshold or the maximum number of simulation rounds, the simulation terminates, outputs a list of consensus-based data usage schemes and their value ranges, and records the complete negotiation process and supporting evidence, forming a traceable "negotiation-evaluation" evidence chain.
4. The method according to claim 1, characterized in that, The specific implementation of the self-reflective reinforcement learning training includes: During the Rollout phase, for the same input question, the model first generates an initial response y1, and then forcibly appends a reflection trigger prompt "To ensure correctness, I'd better verify the previous logic again." to the end of y1, driving the model to continue generating a reflection response y2, forming a complete response sequence [y1 + trigger + y2]. The forced alignment group relative strategy optimization algorithm FA-GRPO is adopted. Based on the algorithm, the intra-group relative advantage of N sequences is generated by the calculation model without the need for an additional value network. A dynamic reward mechanism is introduced into the reward function, including: applying a positive reward to the behavior of successfully correcting the initial wrong answer to the correct answer through reflection, and applying a penalty to the behavior of reflecting the originally correct answer as wrong or reflecting the number of reflection rounds exceeding a preset limit, thereby guiding the model to learn the ability of prudent reasoning and suppressing illusions and logical fallacies.
5. A data value assessment system based on a knowledge mining large model and a simulated intelligent agent for implementing the method as described in any one of claims 1 to 4, characterized in that, Includes the following modules: The knowledge mining module is used to search for domain knowledge and information needed for value assessment and quantification, and to mine and stimulate the knowledge contained within the large model. The value assessment framework module is used to configure and manage common value dimensions preset by experts, providing the system with a structured value assessment perspective; The self-reflective reinforcement learning model building module is used to perform FA-GRPO reinforcement learning training on the basic large language model, integrates reflection triggers and dynamic reward mechanisms, and generates a dedicated large evaluation model with logical self-verification capabilities. The simulation agent module is used to construct supplier, demander and third-party agents based on the evaluation-specific large model, and to simulate the value consensus formation process under multiple rounds of game. The report generation module is used to summarize the list of consensus-based data usage solutions and generate a structured data value assessment report.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the data value assessment method based on a knowledge mining large model and a simulated intelligent agent as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Complex problem reasoning method based on dynamic collaboration of large language model and domain knowledge base
CN118798368A
Vertical domain document question and answer method and system based on knowledge graph enhanced large model
CN119646026A
Question and answer model training method oriented to specific field and intelligent question and answer method and device
CN120654813A
Intelligent data value mining method and system, storage medium and computer program product
CN120873447A
Content auditing method and device, equipment, storage medium and computer program product
CN120974141A
Cited By
Large-Model Auxiliary Logic Function Simulation Verification Method and System for Hardware Design
CN122311003A