Method for realizing rapid retrieval of market main agent based on large model
Through the market entity agent based on large models, dynamic data adaptation, semantic understanding and multi-source heterogeneous data fusion problems in market entity information retrieval are solved, and fast, accurate and efficient market entity information retrieval is achieved, improving the system's response speed and data utilization rate.
Patent Information
- Application Number
- CN202510444981.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has problems such as difficulty in adapting dynamic data, insufficient semantic understanding ability, difficulty in fusion of multi-source heterogeneous data, and unfriendly results in market entities' information retrieval, resulting in long response time, low query efficiency, high maintenance cost and low data utilization rate.
A market entity agent based on large models is adopted to achieve dynamic data adaptation by building NL2SQL capabilities, and semantic-structured transformation and multi-source heterogeneous data fusion are combined with a dual-branch query mechanism to generate intelligent results.
It realizes fast response time (less than 2 seconds), high query efficiency (greater than 85%), low maintenance costs and high data utilization (more than 90%), improving the integrity and accuracy of user experience and data processing.
Smart Images

Figure CN120336324A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and data retrieval, and more particularly, to a method for realizing fast retrieval of market entity agents based on large models. Background Art
[0002] The fast retrieval and accurate analysis of market entities (including enterprises, individual industrial and commercial households, etc.) are the core requirements in scenarios such as supervision, business decision-making, and investment attraction. Traditional retrieval systems rely on structured databases (such as MySQL, Oracle) and fixed rule matching (such as enterprise names, unified social credit codes). However, with the explosion of data volume, the heterogeneity of data sources (i.e., differences in the structure, format, and access methods of data sources), and the complexity of user requirements (such as retrieving "enterprises above a certain scale registered in the past three years and with new energy qualifications"), the existing technologies face many bottlenecks. In the current e-government informatization field, the existing technical implementation architectures have the following disadvantages:
[0003] 1) Static code logic is solidified: The existing system realizes label matching through hard-coded methods (such as if-else rules or fixed SQL templates). When the data table structure changes (such as adding a "carbon emission index" field) or the label system is updated, it is necessary to re-develop the code, with a long response cycle and high maintenance costs.
[0004] 2) Lack of semantic understanding ability: Natural language questions raised by users (such as "finding intelligent manufacturing enterprises registered in the past three years and having high-tech certifications") cannot be directly mapped to database fields and need to be manually disassembled into multi-condition combinations, resulting in low query efficiency.
[0005] 3) Poor adaptability to dynamic data: When the market entity metadata (such as the department data table structure) and the label system (such as industry category labels) change frequently, the traditional system needs to stop the machine to update the code and cannot adapt to the dynamic knowledge base in real time.
[0006] 4) The result presentation is not user-friendly: The returned results are mostly raw data tables, lacking natural language explanations and the context integration of associated labels, resulting in a poor user experience.
[0007] Therefore, there is an urgent need to provide a fast retrieval method that can achieve dynamic data adaptation, semantic structured query conversion, multi-source heterogeneous data fusion, and generate intelligent results. Summary of the Invention
[0008] In view of this, the present invention provides a method for realizing fast retrieval of market entity agents based on large models to solve the following technical problems:
[0009] 1. Dynamic data adaptation: Solve the system reconstruction problem caused by changes in data source structure (such as new fields) and label system expansion (such as the addition of "specialized, refined and innovative" labels). By building a market entity intelligent body with large-model NL2SQL (Natural Language to SQL, natural language queries are converted into structured SQL statements) capabilities, real-time adaptation without code modification is achieved.
[0010] 2. Semantic-structured query conversion: The user's natural language questions (such as "local listed companies with registered capital exceeding 100 million") are automatically parsed into executable database query languages (such as SQL) through the NL2SQL capability of the big model, breaking through the limitations of traditional keyword matching.
[0011] 3. Multi-source heterogeneous data fusion: Integrate cross-departmental data through traditional government system data middle-end products and dynamically build a unified metadata database and tag library through the capabilities of the data middle-end to solve the problem of data silos.
[0012] 4. Intelligent result generation: After integrating the original data with the label information, a natural language description that conforms to the business scenario is generated through a large model to improve the readability of the results.
[0013] The method for realizing rapid retrieval based on a large model of market subject intelligent agents provided by the present invention comprises:
[0014] Data preparation, including:
[0015] Obtain market subject data of multiple departments, use the unified social credit code and the name of the market subject as a joint primary key, and obtain a market subject database table, wherein the data in the market subject database table includes basic information of the market subject and business attribute fields of different departments;
[0016] Extracting metadata from the market entity database table, the metadata including metadata information of the table and metadata information of the table fields, and sorting and classifying and storing them in the market entity database table, the metadata information of the table including the table name and table description, and the metadata information of the table fields including the field name, field type, field length, field constraint conditions, index information, and association information;
[0017] Automatically label each field of each market entity to obtain label data, which is stored in the market entity label library;
[0018] Building intelligent agents based on a common large model, including:
[0019] The user asks a question, including: in response to the user inputting a query requirement through natural language, standardizing the query requirement to obtain a standardized query requirement;
[0020] Retrieve metadata from the market entity database table and convert it into a vector representation. Retrieve label data from the market entity label library and convert it into a vector representation, and establish a field-vector mapping relationship table;
[0021] Enable the dual-branch query mechanism:
[0022] For the first branch, generate a query statement using a predefined Structured Query Language (SQL) expert prompt template, and call the market entity basic information interface to execute the query;
[0023] For the second branch, generate a query statement using a label-specific prompt template, call the market entity label interface to obtain label data, and convert the discrete labels into enterprise standard display data;
[0024] Merge the results. Align the basic information and label data using the unified social credit code as the unique identifier, and append the labels as new fields to the basic information records to obtain the merged data;
[0025] Format and output the merged data;
[0026] The API calls the intelligent agent to achieve dynamic retrieval of market entity information.
[0027] Optionally, retrieve metadata from the market entity database table through a metadata interface.
[0028] Optionally, retrieve label data from the market entity label library through a label interface.
[0029] Optionally, the basic information of the market entity includes: unified social credit code, enterprise name, enterprise type, enterprise registration address, and enterprise business address.
[0030] Optionally, the business attribute fields of different departments include:
[0031] Industrial and commercial registration number, enterprise establishment year, taxpayer identification number, business term, and operation status;
[0032] Whether it is a large-scale enterprise, whether it is a sample enterprise below the scale, and operating income;
[0033] Industry type and industry code;
[0034] And / or, market entity investment intention, technological innovation ability, and business environment feedback.
[0035] Optionally, the large model is the DeepSeek model.
[0036] Optionally, the Structured Query Language (SQL) expert prompt template includes field constraints and table association rules.
[0037] Optionally, the label-specific prompt template includes a label hierarchy and association rules.
[0038] Optionally, the standardization process for the query requirements includes removing special characters and unifying the format.
[0039] Compared with the prior art, the method for realizing fast retrieval of market entity agents based on large models provided by the present invention has at least achieved the following beneficial effects:
[0040] In the prior art, the response time for retrieving market entity information takes more than 5 seconds. However, through the dual-branch query mechanism, the fast retrieval method of the present invention enables the system to quickly respond to user needs, with a response time of less than 2 seconds, greatly improving the retrieval efficiency while ensuring the integrity and accuracy of data processing.
[0041] In the prior art, the semantic understanding ability is lacking, and it needs to be manually disassembled into multi-condition combinations, with a query efficiency lower than 50%. The fast retrieval method of the present invention is automatically parsed based on a large model, with a query efficiency greater than or equal to 85%, greatly improving the query efficiency.
[0042] In the prior art, when the data structure changes, code refactoring is required (the refactoring time is 1 to 3 days), resulting in a long response cycle and high maintenance costs. The fast retrieval method of the present invention realizes implementation adaptation, without code refactoring and without manual intervention, improving efficiency while reducing costs.
[0043] In the prior art, due to the heterogeneity of cross-department data sources, the data utilization rate is relatively low, only reaching about 30%. In the present invention, through the standardization of metadata, the cross-department data utilization rate can reach more than 90%.
[0044] Of course, when implementing any product of the present invention, it is not necessarily required to simultaneously achieve all the above-mentioned technical effects.
[0045] Through the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings, other features and advantages of the present invention will become clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0047] Figure 1 is a flowchart of a method for realizing fast retrieval of market entity agents based on large models provided by an embodiment of the present invention;
[0048] Figure 2 is a flowchart of an automated annotation provided by an embodiment of the present invention;
[0049] Figure 3 is the flowchart of constructing an agent based on a general large model provided by an embodiment of the present invention;
[0050] Figure 4 is a bar chart for counting the number of high-tech enterprises in each county and district of City A. Specific Embodiments
[0051] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present invention.
[0052] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present invention or its application or use.
[0053] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said technologies, methods, and devices should be regarded as part of the specification.
[0054] In all the examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.
[0055] It should be noted that: Like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, further discussion thereof is not required in subsequent drawings.
[0056] Embodiment 1
[0057] Combined with Figures 1 to 3 , the present invention provides a method for realizing rapid retrieval of market entity agents based on a large model, including the following steps:
[0058] S1, data preparation;
[0059] S2, constructing an agent based on a general large model;
[0060] S3, calling the agent through the API to realize dynamic retrieval of market entity information.
[0061] The above step S1, data preparation, includes:
[0062] S101, obtaining market entity data of multiple departments, using the unified social credit code and market entity name as a combined primary key to obtain a market entity database table, and the data in the market entity database table includes the basic information of the market entity and business attribute fields of different departments;
[0063] Specifically, in the process of constructing the fact table of market entities in the whole city (the market entity database table), first, based on the market entity data, the unified social credit code and the market entity name are used as the combined primary key to ensure the uniqueness and accuracy of the data. This step lays a solid foundation for subsequent data integration.
[0064] Next, by supplementing the market entity data of multiple other administrative departments, the content of the market entity information table is further enriched.
[0065] These data not only cover the basic information of market entities but also include the unique business attribute fields of each department, making the fact table of market entities (the market entity database table) more comprehensive and detailed. Among them, the basic information includes the unified social credit code, enterprise name, enterprise type, enterprise registration address, enterprise business address, etc., and the unique business attribute fields of each department. For example, some departments involve industrial and commercial registration numbers, the year of enterprise establishment, taxpayer identification numbers, business periods, operation status; some departments involve whether the enterprise is above-scale, whether it is a sample enterprise below-scale, operating income; some departments involve industry types, industry codes; some departments involve the investment intention of market entities, technological innovation capabilities, feedback on the business environment, etc.
[0066] In the process of data integration, the principle of "mainly based on department data" is always adhered to to ensure that the business fields of each department can be updated in a timely manner and accurately reflected. In this way, a large wide table is successfully constructed, which not only improves the integrity and consistency of the data but also provides strong guarantee for subsequent data analysis and decision support.
[0067] S102, Extract metadata from the market entity database table. The metadata includes the metadata information of the table and the metadata information of the table fields, and is sorted and classified and stored in the market entity database table. The metadata information of the table includes the table name and table description, and the metadata information of the table fields includes the field name, field type, field length, constraint conditions of the field, index information, and association information;
[0068] Specifically, through systematic collection and analysis of the metadata information of the table and the metadata information of the table fields read from the market entity library, the unified storage and management of the metadata information of the market entity library can be comprehensively and accurately realized.
[0069] Extract key metadata information such as table names, table descriptions, field names, field types, and field lengths from the database table structure of the market entity library, and perform normalized sorting and classified storage on them. At the same time, combined with additional attributes such as field constraint conditions, index information, and association relationships, further enrich the content of the metadata. By storing this metadata information, not only can the data structure and logical relationships of the market entity library be clearly presented, but also a solid foundation is provided for subsequent data governance, data quality management, and data application development.
[0070] In addition, the storage of metadata information also supports a dynamic update mechanism to ensure real-time synchronization when the database structure changes, thus ensuring the accuracy and timeliness of the metadata.
[0071] S103. Automatically label each field of each market entity to obtain label data and store it in the market entity label library;
[0072] Specifically, for the market entity data aggregated by each department, adopt a parallel and configurable automated task method to flexibly configure each field, realize efficient and accurate label annotation work, and uniformly store the generated labels in the market entity label library.
[0073] Combined with Figure 2 , specifically, the label annotation process can be illustrated by examples:
[0074] 1. Based on enterprise qualification information, label high-tech enterprises; secondly, according to the enterprise scale, label enterprises with more than 500 employees as "large enterprises", and at the same time label enterprises with less than 50 employees as "small and micro enterprises".
[0075] 2. For enterprise types such as state-owned enterprises, central enterprises, private enterprises, and mixed ownership, perform corresponding type label annotation.
[0076] 3. Combine the industry to which the enterprise belongs (such as manufacturing, finance, information technology, etc.) to label industry classification labels, and label regional labels (such as Beijing, Shanghai, Guangdong, etc.) according to the enterprise's place of registration or main place of business.
[0077] 4. For listed enterprises, label the "listed company" label, and for foreign-funded or joint-venture enterprises, label the "foreign-funded enterprise" or "joint-venture enterprise" label.
[0078] 5. In addition to the above enterprise tags, there are also more than 200 tag information such as key enterprises under focus, gazelle enterprises, provincial gazelle enterprises, municipal gazelle enterprises, specialized and sophisticated enterprises, high-tech enterprises, headquarters enterprises, unicorn enterprises, hidden champions, leading enterprises, enterprises above designated size, listed enterprises, state-owned enterprises, first-level municipal state-owned enterprises, enterprises above a certain scale, construction enterprises with qualification grades, key research and development plans, academician workstations, technology-based small and medium-sized enterprises, engineering technology research centers, maker spaces, incubation carriers, postdoctoral research workstations, provincial engineering research centers, national engineering laboratories, national and local joint engineering laboratories, and enterprise technology centers at or above the provincial level.
[0079] Through this series of automated tag annotations, the characteristics of enterprises can be comprehensively and multi-dimensionally described, providing strong support for data analysis and decision-making.
[0080] The above S2 constructs an intelligent agent based on a general large model, which is based on the large model of LLM (Large Language Model) to construct an AI Agent. The general large model adopted in the embodiments of the present invention is the DeepSeek large model, combined with Figure 3 , and the specific steps include:
[0081] S201. The user raises a question, including: in response to the user's input of a query requirement through natural language, standardizing the query requirement to obtain a standardized query requirement;
[0082] In the initial stage, the user inputs a query requirement through natural language, such as "list the technology-based enterprises in City R with a registered capital of more than 10 million".
[0083] The system first standardizes the user input to ensure the accuracy of subsequent processing. In some alternative embodiments, the standardization process includes operations such as removing special characters and unifying formats. This step not only improves the normativity of the query statement but also lays a foundation for subsequent semantic parsing and data processing. Through input standardization, the system can effectively avoid parsing errors caused by differences in user input, thereby improving the overall query efficiency.
[0084] S202. Obtain metadata from the market entity database table and convert it into a vector representation, obtain tag data from the market entity tag library and convert it into a vector representation, and establish a field-vector mapping relationship table;
[0085] Specifically, in some alternative embodiments, the system calls the / metadata (metadata) interface to obtain the basic fields of market entities, such as metadata information such as registered capital and establishment time, and at the same time obtains the tag system such as industry classification and credit rating through the / tags (tags) interface. Subsequently, the system converts this structured data into a vector representation and establishes a field-vector mapping relationship table.
[0086] Vectorization not only improves the computational efficiency of data, but also makes subsequent query and matching operations more flexible and efficient. Through this step, the system can convert complex structured data into a unified vector form, providing data support for the dual-branch query mechanism.
[0087] S203. Activate the dual-branch query mechanism
[0088] It should be noted that in order to improve query efficiency, the system adopts a dual-branch parallel processing mechanism. In the resource allocation stage, the system allocates independent memory space and computing resources to the two threads of basic information processing and label information processing, ensuring that both can run efficiently simultaneously. This parallel processing method not only shortens the overall query time, but also improves the resource utilization rate of the system. Through the dual-branch query mechanism, the system can quickly respond to user needs while ensuring the integrity and accuracy of data processing.
[0089] S2031. For the first branch, use a predefined structured query language expert prompt template to generate a query statement, and call the market entity basic information interface to execute the query;
[0090] Specifically, in the first branch, the system first uses a predefined SQL expert Prompt template to generate a query statement. Combining with vectorized metadata, the system generates an SQL statement similar to SELECT * FROM enterprises WHERE reg_capital > 10000000. The SQL expert Prompt template is the structured query language expert prompt template, which includes field constraints, table association rules, etc. Subsequently, the system calls the / enterprise-info (market entity basic information) interface to execute the query and performs integrity verification on the returned data to ensure that there are no missing fields and no abnormal values. This process ensures the accurate acquisition of basic information and provides a reliable basis for subsequent data fusion.
[0091] The following is a specific embodiment of the SQL expert Prompt template. Of course, specific limitations are imposed on the SQL expert Prompt template here.
[0092] # Role: You are a database expert proficient in SQL language, proficient in PostgreSQL, and also good at interpreting and analyzing data
[0093] # Task: Your task is to understand the user's input and context content, write an SQL query, call the tool to query and obtain the result, and present, interpret, and analyze the query result in combination with the user's question
[0094] # Key steps:
[0095] 1. Identify and judge the content input by the user. If the content involves current events, social issues, and situations that violate morality and laws and regulations, output: "The questions raised exceed the scope that should be answered. Please ask questions related to the company's business, otherwise no answer can be given."
[0096] 2. Based on the content input by the user and the context information, form a content classification. According to the content classification, retrieve the data table structure information from the knowledge base "Data Structure Description" according to the following rules:
[0097] - If the content classification is related to the enterprise, retrieve the "Enterprise Information Table".
[0098] Note: Be sure to strictly obtain the corresponding retrieval keywords according to the above classification, and do not generate new retrieval keywords. If you think the user's question cannot be matched to a suitable classification, please output a prompt: To ensure accurate query information, please describe your needs in more detail.
[0099] 3. Based on the content input by the user and the context information, form a complete question that conforms to the user's intention, and use this as the input to retrieve the SQL statement reference example in the knowledge base "SQL Examples".
[0100] 4. Based on the understanding of the context and the user's question, write an SQL query statement according to the retrieved data table structure information and the SQL reference example. Note that if the content classification does not match the classification in the reference example, this example will be ignored. In addition, not all cases have example references. When there is no example, write the SQL statement according to your own understanding and knowledge.
[0101] 5. Remove the redundant comments, line breaks and other useless information in the SQL statement, and output a pure and directly executable SQL statement.
[0102] 6. Execute the SQL query to obtain the result.
[0103] 7. Read the query result, and combine the historical conversation content to present, interpret and analyze the query result.
[0104] # Precautions when writing SQL:
[0105] 1. Be sure to write the SQL statement according to the data table structure description provided by the context, ensure that only the table names and field names mentioned in the data table structure description are used, and refer to the explanations of the fields.
[0106] 2. Ensure that the SQL is compatible with postgressql10.0
[0107] 3. Only use Simplified Chinese.
[0108] 4. Output only one complete SQL statement without comments, ensuring that it can be directly executed and obtain the expected results
[0109] 5. For fields of string and long text types, unless the user specifically states otherwise, use the LIKE operation instead of the equal operation. For example: WHERE product model LIKE N'%keyword%' instead of WHERE product model = 'keyword'
[0110] 6. Division processing: Refer to the following template to avoid errors:
[0111] CASE WHEN [divisor] = 0 THEN 0
[0112] ELSE CAST([dividend] AS FLOAT) / [divisor]
[0113] END AS [result column name]
[0114] # Requirements for data presentation, interpretation, and analysis:
[0115] 1. All data already meets the conditions in the user's question (such as gender of personnel, date range)
[0116] 2. Directly use the provided data analysis without questioning whether the data meets the conditions
[0117] 3. Do not need to filter or confirm the data category / time range again
[0118] 4. When the data is [] or empty, reply "No relevant data was found", and do not fabricate data
[0119] 5. List the detailed data. Give priority to listing the data in a table. If the data exceeds 10, unless the user's question clearly requires listing all the data, you only need to randomly list 5 representative records. Otherwise, list all the records
[0120] 6. When some records are omitted, an explanation must be made
[0121] 7. Provide an overview and summary of the data, which must include the total number of original records
[0122] 8. Identify trends, anomalies, and provide analysis and suggestions
[0123] 9. Data presentation processing:
[0124] - Number of decimal places: All decimals are rounded to two decimal places
[0125] - Use the comma counting method: 1,234,567
[0126] -For fields such as ratios, proportions, and similar meanings, display them as percentages. For example, display 0.1236 as 12.36%, retaining two decimal places
[0127] -The display format for dates is: YYYY-MM-DD. For example: 2024-08-01
[0128] -Ensure that the correct markdown syntax is used, especially for headings and tables
[0129] #Other considerations
[0130] 1. Do not output the intermediate thinking process, only output the final result
[0131] 2. The most important point is to output as sql!!! Do not output comments and explanatory texts, do not contain special characters, do not contain the character '\n' and line breaks, and Chinese aliases are not allowed in sql. Only output pure sql statements
[0132] In the above SQL expert Prompt template, PostgreSQL is an open-source, extensible relational database management system that can be freely used, modified, and distributed. It supports various data types, including JSON, XML, arrays, etc., and supports custom data types, functions, triggers, etc.
[0133] The LIKE operation is a commonly used fuzzy matching method. LIKE is a fuzzy matching operator in SQL, used to search for strings that match a specific pattern in the WHERE clause. It allows the use of wildcards for flexible pattern matching.
[0134] S2032, for the second branch, use the label-specific prompt word template to generate a query statement, call the market entity label interface to obtain label data, and convert the discrete labels into enterprise standard display data;
[0135] In the second branch, the system uses the label-specific Prompt template to generate a query statement, such as SELECT * FROM enterprise_tags WHERE tag_type = 'industry'. Then, the system calls the / enterprise-tags interface to obtain label data and converts the discrete labels into enterprise standard display data. For example: convert {Enterprise A: {"Industrial enterprise"}, Enterprise B: {"Industrial enterprise"}} into: "data": [{"name": "Enterprise A", "isIndustry": "yes"}, {"name": "Enterprise B", "isIndustry": "yes"}]
[0136] This normalization process makes the labeled data easier to understand and analyze, facilitating subsequent result merging and visualization.
[0137] In some alternative embodiments, the label-specific Prompt template includes label hierarchy relationships and association rules.
[0138] The label-specific Prompt template is as follows:
[0139] # Role
[0140] You are an experienced and highly professional senior SQL expert with profound expertise in dynamically generating SQL query statements based on natural language descriptions and the table structures in the label knowledge base. You are particularly good at handling complex multi-table association queries and always adhere to the principle of only generating query statements (SELECT), resolutely avoiding generating any data manipulation statements. The core task is to accurately query the corresponding list of market entities or the corresponding quantity according to the market entity label information. During the label generation process, it is necessary to associate with the basic information table of the market entity. Additionally, currently during the label association process, first perform explicit matching on the labels and temporarily do not consider the issue of label hierarchy.
[0141] ## Skills
[0142] Skill 1: Generate SQL query statements
[0143] 1. Carefully receive the natural language description provided by the user and precisely identify the data they want to query.
[0144] 2. Dynamically generate SQL query statements that meet the requirements based on the user's description and the table structures in the knowledge base. When generating the statements, ensure that the JOIN statements are accurate to correctly handle the association relationships between multiple tables.
[0145] ## Limitations:
[0146] - Only allow generating SELECT statements and absolutely prohibit generating any data manipulation statements.
[0147] - Must be able to automatically identify and correctly handle the association relationships between multiple tables and accurately generate JOIN statements.
[0148] - The generated SQL query statements must be based on the table structure information in the knowledge base to ensure accuracy and reasonableness.
[0149] In the above label-specific Prompt template, resolutely avoid generating any data manipulation statements, such as INSERT for inserting new data into a table, DELETE for deleting existing data in a table, and UPDATE for updating data in a table.
[0150] S204, Result merging: Align the basic information and label data using the unified social credit code as the unique identifier, and append the labels as new fields to the basic information records to obtain the merged data;
[0151] Specifically, in the result merging stage, the system aligns the basic information and label data using unique identifiers such as the unified social credit code. Subsequently, the system appends the labels as new fields to the basic information records. This data fusion method not only expands the data dimension but also improves the readability and analytical value of the data. Through conflict handling, the system ensures the integrity and consistency of the merged data.
[0152] S205, Format the output of the merged data;
[0153] Specifically, the system formats the output of the merged data. In the structured transformation stage, the system automatically identifies numerical fields (such as registered capital) and performs thousand - digit formatting, while converting classification labels into readable icons or color markers. In addition, the system uses a preset template to generate summary descriptions and dynamically generates data insights (such as "Qualified enterprises account for 12% of the total number"). This natural language generation function makes the output results more intuitive and user - friendly, helping users quickly understand the insights behind the data. Through formatted output, the system not only improves the readability of the data but also enhances the user experience.
[0154] Subsequently, code calls and implementations are carried out in the business system. It is necessary to perform API calls of the AIAgent (intelligent agent) in the business system to implement the dynamic retrieval function of market entity information.
[0155] The present invention adopts a dynamic metadata - driven mechanism, guiding the large model to generate SQL by real - time reading of metadata to achieve automatic adaptation to data structure changes.
[0156] The present invention supports high concurrency and dynamic expansion through dual - branch hybrid retrieval, separating the queries for basic information and label information.
[0157] The present invention adopts domain - customized Prompt engineering. By defining the Prompt for the SQL expert role, it constrains the large model to output query statements that conform to the database specifications.
[0158] Example 2
[0159] S1, Data preparation, including:
[0160] S101, Obtain the market entity data of multiple departments, using the unified social credit code and the market entity name as the composite primary key to obtain the market entity database table. The data in the market entity database table includes the basic information of the market entity and the business attribute fields of different departments;
[0161] Since there is a large amount of data of market entities (enterprises), it is not shown here.
[0162] S102. Extract metadata from the market entity database table. The metadata includes the metadata information of the table and the metadata information of the table fields, and is sorted, classified and stored in the market entity database table. The metadata information of the table includes the table name and the table description. The metadata information of the table fields includes the field name, the field type, the field length, the constraint conditions of the field, the index information and the association information.
[0163] Refer to Table 1 below. Table 1 is an example of the enterprise basic information table.
[0164] Table 1 Example of Enterprise Basic Information Table
[0165] Field Name Field Comment Field Type Field Length Id Enterprise ID VARCHAR 36 enterprise_name Enterprise Name VARCHAR 255 credit_code Unified Social Credit Code VARCHAR 30 address Registered Address Text 0 legal_person Legal Representative VARCHAR 255 operate_adress Business Address Text 0 qylxr Enterprise Contact VARCHAR 255 lxfs Contact Information VARCHAR 11 ....
[0166] Extract metadata from the market entity database table. The metadata includes the metadata information of the table and the metadata information of the table fields, such as enterprise ID, enterprise name, unified social credit code, registered address, legal representative, business address, enterprise contact person, contact information, etc. Refer to Table 2 below. Table 2 is an example of the enterprise basic data.
[0167] Table 2 Example of Enterprise Basic Data
[0168]
[0169] S103. Automatically annotate each field of each market entity to obtain tag data, which is stored in the market entity tag library; the tag data may include: tag ID, tag name, tag type, classification directory ID, tag source, tag active number, entity member unique code, etc. Refer to Table 3 below. Table 3 is an example of the tag basic information table.
[0170] Table 3 Example of Tag Basic Information Table
[0171]
[0172]
[0173] Refer to Table 4 below. Table 4 is an example of the tag basic data
[0174] Table 4 Example of Tag Basic Data
[0175]
[0176] Next is to build an intelligent agent based on the general large model. For the specific steps, refer to S201 - S205 in Embodiment 1.
[0177] S201. The user raises a question, including: in response to the user's input of a query requirement through natural language, standardizing the query requirement to obtain a standardized query requirement;
[0178] For example, the user raises a question: the number of high-tech enterprises in each region of City A.
[0179] S202. Obtain metadata from the market entity database table and convert it into a vector representation, obtain label data from the market entity tag library and convert it into a vector representation, and establish a field-vector mapping relationship table;
[0180] At this time, it is necessary to establish a sub-field-vector mapping relationship table for the enterprises with the label of high-tech enterprises in the label data and the enterprises in the market entity database table.
[0181] S203. Activate the double-branch query mechanism, use the predefined structured query language expert prompt word template to generate a query statement to call the market entity basic information interface to execute the query, use the label-specific prompt word template to generate a query statement, call the market entity tag interface to obtain label data, and convert the discrete labels into enterprise standard display data;
[0182] S204. Result merging. Align the basic information and label data with the unified social credit code as the unique identifier, append the label as a new field to the basic information record to obtain the merged data;
[0183] S205. Format and output the merged data, and the output format can be various forms, such as charts or tables, etc.
[0184] Finally, call the intelligent agent through the API to achieve dynamic retrieval of market entity information.
[0185] For example, input in the intelligent agent: Please count the number of high-tech enterprises in each region of City A. The intelligent agent can count the number of high-tech enterprises in each county and district of City A in the form of a table or a chart. Table 5 is the statistical table of the number of high-tech enterprises in each county and district of City A. Figure 4 It is the bar chart of the statistical table of the number of high-tech enterprises in each county and district of City A.
[0186] Table 5 Statistical Table of the Number of High-Tech Enterprises in Each County and District of City A
[0187] Serial Number Region Number of High-Tech Enterprises 1 Area A 1116 2 Area B 727 3 Area C 352 4 Area D 345 5 Area E 293 6 Area F 261 7 Area G 221 8 Area H 225 9 Area I 178 10 Area J 172 11 Area K 104 12 Area L 8 13 County M 72 14 County N 44 15 Area O 39
[0188] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should understand that the above embodiments can be modified without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A method for an intelligent agent of market entities based on a large model to achieve rapid retrieval, characterized in that, Including: Data preparation, including: Obtaining market entity data of multiple departments, using the unified social credit code and market entity name as the combined primary key to obtain the market entity database table, where the data in the market entity database table includes the basic information of the market entity and the business attribute fields of different departments; Extracting metadata from the market entity database table, where the metadata includes table metadata information and table field metadata information, and organizing and classifying and storing them in the market entity database table. The table metadata information includes the table name and table description, and the table field metadata information includes the field name, field type, field length, field constraint conditions, index information, and association information; Automatically annotating each field of each market entity to obtain label data and storing it in the market entity label library; Building an intelligent agent based on a general large model, including: The user poses a question, including: in response to the user's input of a query requirement in natural language, standardizing the query requirement to obtain a standardized query requirement; Obtaining metadata from the market entity database table and converting it into a vector representation, obtaining label data from the market entity label library and converting it into a vector representation, and establishing a field-vector mapping relationship table; Enabling a dual-branch query mechanism: For the first branch, generating a query statement using a predefined structured query language expert prompt template and calling the market entity basic information interface to execute the query; For the second branch, generating a query statement using a label-specific prompt template, calling the market entity label interface to obtain label data, and converting the discrete labels into enterprise standard display data; Result merging, aligning the basic information and label data using the unified social credit code as the unique identifier, appending the label as a new field to the basic information record to obtain the merged data; Formatting and outputting the merged data; Calling the intelligent agent through the API to achieve dynamic retrieval of market entity information.
2. The method for realizing fast retrieval of market entity agents based on large models according to claim 1, wherein, Obtaining metadata from the market entity database table through the metadata interface.
3. The method for realizing fast retrieval of market entity agents based on large models according to claim 1, characterized in that, Obtaining label data from the market entity label library through the label interface.
4. The method for realizing fast retrieval of market entity agents based on large models according to claim 1, characterized in that, The basic information of the market entity includes: unified social credit code, enterprise name, enterprise type, enterprise registration address, and enterprise business address.
5. The method for realizing fast retrieval of market entity agents based on large models according to claim 1, wherein The business attribute fields of different departments include: Industrial and commercial registration number, enterprise establishment year, taxpayer identification number, business term, and operation status; Whether it is a scale-up enterprise, whether it is a sample enterprise below the scale, and operating income; Industry type and industry code; And / or, market entity investment intention, technological innovation ability, and business environment feedback.
6. The method for realizing fast retrieval of market entity agents based on large models according to claim 1, wherein The large model is the DeepSeek model.
7. The method for achieving fast retrieval of market entity agents based on large models according to claim 1, characterized in that, The structured query language expert prompt template includes field constraints and table association rules.
8. The method for realizing fast retrieval of market entity agents based on large models according to claim 1, characterized in that The label-specific prompt template includes label hierarchy relationships and association rules.
9. The method for realizing fast retrieval of market entity agents based on large models according to claim 1, wherein The standardizing the query requirement includes removing special characters and unifying the format.
Citation Information
Cited By
Intelligent question and answer method and system, medium, equipment and program product
CN121233705A