Large model-based data acquisition and processing configuration generation method and system
By using large models for intent recognition and multi-round interaction optimization, and automatically generating data processing configuration scripts, the problems of low efficiency, high professional skill requirements, and communication barriers in traditional data processing are solved, achieving efficient and accurate data processing and personalized services.
Patent Information
- Application Number
- CN202510757704.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional data processing methods are inefficient, error-prone, require specialized programming skills, and cannot meet diverse user needs, creating a communication barrier between user requirements and technology.
A data acquisition and processing configuration generation method based on a large model is adopted. Through intent recognition and semantic analysis of user needs, multiple rounds of interaction optimization are carried out to automatically generate configuration scripts that conform to the user platform specifications.
Improve data processing efficiency and accuracy, lower barriers to entry, meet diverse user needs, and enhance user experience.
Smart Images

Figure CN120910068A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of big data processing, and particularly relates to a data collection and processing configuration generation method and system based on a large model. BACKGROUND
[0002] In the technical field of big data processing, data collection, processing and storage are key links. Traditional data processing methods usually perform filtering, aggregation, conversion and storage operations on data based on structured query language (SQL). In the field of artificial intelligence technology, natural language processing (NLP) technology has been widely applied to various scenarios such as speech recognition, machine translation, sentiment analysis, etc. Through deep learning algorithms and large-scale corpus training, NLP technology can understand and generate natural language, realizing efficient interaction between humans and machines.
[0003] The traditional data processing method has the following problems:
[0004] (1) The problem of low efficiency and error-prone manual SQL statement writing. Traditional data processing methods require users to manually write SQL statements to implement data filtering, aggregation, conversion and storage operations, which is inefficient and prone to errors, and requires professional operation.
[0005] (2) The problem that existing data processing tools and frameworks require users to have certain programming skills. Existing data processing tools and frameworks, such as Apache Hadoop, Apache Spark, etc., provide more efficient data processing capabilities, but these tools and frameworks also require users to have certain programming skills, with a high threshold.
[0006] (3) The problem that existing solutions cannot meet users' diverse data processing needs and lack flexibility. Existing data processing methods and tools cannot meet users' diverse data processing needs and lack flexibility, and cannot be customized according to users' actual needs.
[0007] (4) The problem of communication barriers between user needs and data processing technology. Traditional data processing methods and tools cannot effectively communicate with users and accurately understand users' needs, resulting in inaccurate or unsatisfactory data processing results.
[0008] Therefore, how to improve the efficiency and accuracy of data processing, reduce the threshold, and meet users' diverse needs is a problem to be solved in the current technical field. SUMMARY
[0009] To solve the above problems existing in the traditional data processing mode, the application provides a data acquisition and processing configuration generation method and system based on a large model, which greatly improves the efficiency and accuracy of data processing, reduces the threshold of data processing, and meets the diversified data processing needs of different users.
[0010] To achieve the above purpose, the application adopts the following technical solutions:
[0011] In an embodiment of the application, a data acquisition and processing configuration generation method based on a large model is provided, which comprises:
[0012] The user inputs a natural language requirement, the large model is used for intent recognition and semantic analysis of the input text, the user's requirement is preliminarily identified, including task type and key parameter;
[0013] The system proposes a detailed question, dynamically updates the requirement and provides real-time suggestions, and interacts with the user in multiple rounds to further improve the user's requirement and obtain more detailed information, helping the system accurately identify the requirement;
[0014] After multiple rounds of interaction, the system will summarize all known information and confirm the final requirement with the user, and if the user has any objection to the requirement description or wishes to modify it, the system will adjust again until the requirement is completely confirmed;
[0015] Once the requirement is confirmed, the system will automatically generate a configuration script in a format that meets the user's target platform specifications according to the identified task type and details.
[0016] Further, the method further comprises: built-in prompt word template, training large model.
[0017] Further, the built-in prompt word template and the training of the large model comprise:
[0018] The built-in prompt word template and the training of the large model improve the intent recognition ability;
[0019] Define the mapping relationship between the template key field and the data processing capability.
[0020] Further, the mapping relationship between the template key field and the data processing capability comprises:
[0021] Define the key field commonly used in data processing tasks;
[0022] Define a data processing capability library, which is a set of function modules or components for processing user requirements; each component is associated with a specific task type and key field;
[0023] Map the above key field and data processing capability library component.
[0024] In an embodiment of the present application, a large model-based data acquisition and processing configuration generation system is also provided, which comprises:
[0025] An initial requirement capturing module is configured to input natural language requirements by a user, perform intent recognition and semantic analysis on the input text by using a large model, and preliminarily identify the requirements of the user, including task types and key parameters.
[0026] A multi-round interaction optimization module is configured to propose refined questions, dynamically update requirements, and provide real-time suggestions by the system, and perform multi-round interactions with the user, further refine the requirements of the user, obtain more detailed information, and help the system accurately identify the requirements.
[0027] An intent verification and confirmation module is configured to, after the multi-round interactions, aggregate all known information by the system, and confirm the final requirements with the user, and if the user has any objection to the requirement description or wishes to make modifications, the system will adjust again until the requirements are completely confirmed.
[0028] A configuration script generation module is configured to, once the requirements are confirmed, automatically generate a configuration script in a format that conforms to the specification of the target platform of the user according to the identified task types and details.
[0029] Further, the method further comprises a large model training module configured to embed a prompt word template and train the large model.
[0030] Further, the large model training module is specifically configured to:
[0031] Embed a prompt word template and train the intent recognition capability of the large model.
[0032] Define a mapping relationship between template key fields and data processing capabilities.
[0033] Further, defining the mapping relationship between the template key fields and the data processing capabilities comprises:
[0034] Defining common key fields in data processing tasks;
[0035] Defining a data processing capability library, which is a set of functional modules or components for processing user requirements; each component is associated with a specific task type and key field.
[0036] Mapping the above key fields and data processing capability library components.
[0037] In an embodiment of the present application, a computer device is also provided, which comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements the large model-based data acquisition and processing configuration generation when executing the computer program.
[0038] In an embodiment of the present application, a computer readable storage medium is also provided, which stores a computer program for generating a large model-based data collection and processing configuration.
[0039] Advantages:
[0040] 1. Improve data processing efficiency and accuracy: The present application uses a large model to identify and refine the problem, automatically generates an execution script, reduces the step of manually writing SQL statements, and improves the efficiency of data processing. At the same time, through multi-round interaction optimization and demand confirmation, the accuracy of the demand is ensured, and the accuracy of the data processing is improved.
[0041] 2. Reduce the threshold and meet the diverse needs of users: The present application does not require users to have professional programming skills, but only needs to interact with the system through natural language, which can complete complex data processing tasks. At the same time, through dynamic updating of demand and real-time suggestions, the system can provide personalized suggestions and operation options according to the actual situation of the user, and meet the diverse needs of the user.
[0042] 3. Improve user experience: The present application guides and interacts with the user through feedback and multi-round interaction, understands the real needs of the user, provides personalized services and operation suggestions, and improves the user experience. At the same time, through demand confirmation and modification, the system can ensure that the user's demand is accurately understood and implemented, further improving the user experience. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 is a flow chart of the large model-based data collection and processing configuration generation method of the present application;
[0044] Figure 2 is a structural schematic diagram of the large model-based data collection and processing configuration generation system of the present application;
[0045] Figure 3 is a structural schematic diagram of the computer equipment of the present application. DETAILED DESCRIPTION
[0046] The principles and spirits of the present application will be described below with reference to a number of exemplary embodiments, and it should be understood that these embodiments are given only to enable those skilled in the art to better understand and implement the present application, and not to limit the scope of the present application in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0047] Those skilled in the art understand that the embodiments of the present application can be implemented as a system, a system, a device, a method or a computer program product. Therefore, the present disclosure can be embodied in the form of a complete hardware, a complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0048] According to the embodiments of the present application, a large model-based data acquisition and processing configuration generation method is proposed, which greatly improves the efficiency and accuracy of data processing, reduces the threshold of data processing, and meets the diversified data processing needs of different users.
[0049] The principles and spirits of the present application will be explained in detail below with reference to several representative embodiments of the present application.
[0050] Figure 1 It is a large model-based data acquisition and processing configuration generation method flow chart of the present application. As shown in Figure 1 The implementation steps of the method are as follows:
[0051] S1, built-in prompt word template, training large model.
[0052] Built-in prompt word template, training large model's intention recognition ability;
[0053] Define the mapping relationship between template key fields and data processing capabilities.
[0054] S2, initial demand capture.
[0055] Adopting natural language processing technology and big data analysis technology, through the user input natural language demand, using large model to input text intention recognition and semantic analysis, preliminary identification of user's demand, including data acquisition, data processing, data storage and other key task type and related parameters.
[0056] S3, multi-round interaction optimization.
[0057] Adopting multi-round dialogue technology, through the system puts forward the detailed question, dynamic update demand and real-time suggestion and other ways, with the user for multi-round interaction, further perfect the user demand, get more detail information, help the system accurate identification demand.
[0058] S4, intention verification and confirmation.
[0059] Adopting demand confirmation and modification technology, after multi-round interaction, the system will summarize all the known information, and confirm the final demand to the user, if the user has any objection to the demand description or hope to modify, the system will adjust again, until the demand is completely confirmed.
[0060] S5, configuration script generation.
[0061] Using script generation technology, once the requirements are confirmed, the system will automatically generate configuration scripts according to the identified task type and details. These configuration scripts will be automatically generated according to different task scenarios, and the format will conform to the specifications of the user's target platform (such as MySQL, ClickHouse, etc.).
[0062] Large model core function: responsible for semantic understanding and interaction logic in S2 (initial requirement capture), S3 (multi-round interaction optimization), and S4 (intention verification and confirmation).
[0063] System auxiliary function: S5 (configuration script generation) is automatically generated by the system according to the mapping relationship, and the large model provides support for key parameter extraction.
[0064] Using user interface technology, through a friendly user interface, users can easily input natural language descriptions, view requirement confirmation and modification results, download generated script files, etc., improving user experience.
[0065] It should be noted that although the operations of the method of the present application are described in a specific order in the above embodiments and drawings, this does not require or imply that the operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired results. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step, and / or one step can be divided into multiple steps.
[0066] In order to more clearly explain the above-mentioned large model-based data collection and processing configuration generation method, a specific embodiment will be described below, however, it should be noted that this embodiment is only to better illustrate the present application, and does not constitute an improper limitation on the present application.
[0067] Example 1:
[0068] This embodiment mainly describes the specific implementation process of a large model-based data collection and processing configuration generation method.
[0069] S1: Built-in prompt word template, train large model.
[0070] In a preferred embodiment, step S1 includes:
[0071] S11, built-in prompt templates for training the intent recognition capabilities of large models. These templates include common task requirements and operation instructions in the data collection and processing process. These templates are structured prompt templates, used to guide the large model to identify key fields (such as data source, aggregation method) in user requirements and associate them with pre-set data processing capabilities. For example, when the user mentions "daily aggregation", the template will guide the model to extract "time granularity = day" and "operation type = aggregation". Through these templates, large models (such as GPT-4, ChatGLM, DeepSeek-R1, etc.) can understand and extract potential operation intentions and key information in conversations with users, helping the system automatically identify and respond to different data processing needs. Template examples are as follows:
[0072] ① Data collection task template
[0073] Task description: Collect data from [data source] and store it in [target storage].
[0074] Input information:
[0075] - Data source type: [ClickHouse / MySQL / Elasticsearch, etc.];
[0076] - Data source connection information: [host address, port, database name, table name, etc.];
[0077] - Fields to be collected: [field list];
[0078] - Time range for collection: [start time] to [end time];
[0079] - Target storage type: [ClickHouse / MySQL / Elasticsearch, etc.];
[0080] - Target storage table name: [table name];
[0081] - Storage method: [full synchronization / incremental synchronization];
[0082] - Storage format: [CSV / Parquet / JSON, etc.]。
[0083] ② Data cleaning and filtering task template
[0084] 1) Data cleaning (de-duplication, missing value filling, etc.)
[0085] Task description: Clean the collected data, remove duplicates and fill in missing values. Input information:
[0086] - Data source type: [ClickHouse / MySQL, etc.];
[0087] - Data source connection information: [database address, table name, etc.] ;
[0088] - Cleaning rules: [de-duplication, missing value filling, outlier removal, etc.] ;
[0089] - Fields that need to be cleaned: [field list] ;
[0090] - Outlier definition: [condition to define outliers] ;
[0091] - Missing value handling method: [mean filling, median filling, specified value filling].
[0092] 2) Data filtering (filtering data by conditions)
[0093] Task description: Filter data according to given conditions.
[0094] Input information:
[0095] - Data source type: [ClickHouse / MySQL, etc.] ;
[0096] - Data source connection information: [database address, table name, etc.] ;
[0097] - Fields that need to be filtered: [field list] ;
[0098] - Filtering conditions: [field name] [operator] [value] (e.g., `amount>1000`) ;
[0099] - Target storage type after filtering: [ClickHouse / MySQL, etc.] ;
[0100] - Target table name: [table name].
[0101] ③ Data type conversion template
[0102] Task description: Convert fields in a data table from one type to another.
[0103] Input information:
[0104] - Data source type: [ClickHouse / MySQL, etc.] ;
[0105] - Data source connection information: [database address, table name, etc.] ;
[0106] - Fields that need to be converted: [field list] ;
[0107] - Source type: [source field type] ;
[0108] - Target type: [target field type].
[0109] 4. Data aggregation task template
[0110] 1) Aggregation based on time dimension
[0111] Task description: Aggregate data by [time field] with [aggregation method] and store the result in [target storage].
[0112] Input information:
[0113] - Data source type: [ClickHouse / MySQL, etc.];
[0114] - Data source connection information: [database address, table name, etc.];
[0115] - Aggregation fields: [field list];
[0116] - Aggregation method: [sum / average / max / min, etc.];
[0117] - Aggregation granularity: [day / month / year, etc.];
[0118] - Target storage type: [ClickHouse / MySQL, etc.];
[0119] - Target table name: [table name];
[0120] - Storage format: [CSV / Parquet, etc.].
[0121] 2) Aggregation based on other dimensions (such as device, region, etc.)
[0122] Task description: Aggregate data by [dimension field] with [aggregation method] and store the result in [target storage].
[0123] Input information:
[0124] - Data source type: [ClickHouse / MySQL, etc.];
[0125] - Data source connection information: [database address, table name, etc.];
[0126] - Aggregation fields: [field list];
[0127] - Aggregation method: [sum / average / max / min, etc.];
[0128] - Aggregation dimension: [device name / region / user, etc.];
[0129] - Target storage type: [ClickHouse / MySQL, etc.];
[0130] - Target table name: [table name];
[0131] - Storage format: [CSV / Parquet, etc.]
[0132] The templates for data collection and processing can be iteratively expanded according to actual business needs to meet the requirements of different business scenarios in the production environment.
[0133] S12. Define the mapping relationship between key template fields and data processing capabilities. Through this mapping relationship, the system can accurately identify the corresponding processing capabilities and generate configuration scripts based on user needs and key fields in natural language input.
[0134] ① Definition of key fields
[0135] First, it's necessary to define key fields commonly encountered in data processing tasks. These fields typically appear in the user's natural language input and help the large model identify the specific requirements of the task. Common key fields include:
[0136] Data source: Indicates the source of the data, such as "MySQL", "Elasticsearch", "Hadoop", etc.
[0137] Target platform: Indicates the target storage or processing platform for data processing, such as "ClickHouse", "MySQL", "Kafka", etc.
[0138] Fields: Indicate the data fields that need to be processed, such as "sales data", "inflow rate", "transaction amount", etc.
[0139] Operation type: Indicates the operation performed on the data, such as "aggregate", "clean", "transform", "filter", etc.
[0140] Aggregation method: Indicates how the data is aggregated, such as "sum", "count", "average", etc.
[0141] Time granularity: Indicates the time dimension when aggregating data, such as "by day", "by month", "by year", etc.
[0142] Filter criteria: Indicate the data filtering conditions, such as "sales amount greater than 1000" or "status is successful".
[0143] Storage table: Indicates the target table name for data storage, such as "sales_data", "transaction_records", etc.
[0144] ② Definition of Data Processing Capability Library
[0145] A data processing capability library refers to a set of functional modules or components that can handle user needs. Each component is associated with a specific operation type and key fields. Common capability library components include:
[0146] Data collection component: Collects data from a specified data source. Corresponding key fields: data source, field.
[0147] Data cleaning component: Formats, removes outliers, fills missing values, and other operations on data. Corresponding key fields: field, filtering condition.
[0148] Data conversion component: Converts data from one format to another. For example, CSV to JSON, JSON to table, etc. Corresponding key fields: field, target platform.
[0149] Data aggregation component: Aggregates data, such as sum, count, average, etc. Corresponding key fields: aggregation method, field, time granularity.
[0150] Data storage component: Stores processed data to a target platform (such as ClickHouse, MySQL, etc.). Corresponding key fields: target platform, storage table.
[0151] Data filtering component: Filters data according to specified conditions. Corresponding key fields: filtering condition, field.
[0152] ③ Definition of mapping relationship
[0153] By mapping the above key fields with the component library, the large model can select appropriate components according to the key fields in the natural language input, and generate configuration scripts according to these fields. Table 1 below is a simple mapping relationship example:
[0154] Table 1
[0155]
[0156] Based on the above field and capability mapping relationship, the system can automatically parse and match to appropriate component library components according to the user input natural language requirements, thereby generating configuration scripts.
[0157] S2: Initial requirement capture.
[0158] In a preferred embodiment, step S2 includes:
[0159] S21, receiving user input.
[0160] The user inputs natural language description on the system interface, describing the requirements for data collection, data processing, data storage, etc. For example, the user inputs "I want to collect data from the users table in MySQL, aggregate by day, and store the results in the daily_users table in ClickHouse, with the aggregation method being sum."
[0161] S22, intent recognition and semantic analysis.
[0162] The system employs natural language processing techniques and big data analysis techniques to perform intent recognition and semantic analysis on the user's natural language input, and initially identifies the user's requirements. For example, the system identifies that the user needs to perform data collection, data processing, and data storage operations, and needs to aggregate daily, with the aggregation method being summation.
[0163] S23, feedback guidance.
[0164] The system provides feedback to the user on the initial understanding, and guides the user based on ambiguous or unclear requirements, asking further questions to clarify the details of the requirements.
[0165] The system handles ambiguous or unclear requirements input by the user through the following mechanisms:
[0166] Triggering condition: When the large model detects that key parameters are missing in the user's requirements (such as not specifying the time range of the data source, the aggregation method is not clear) or there are logical contradictions, the follow-up process is automatically triggered.
[0167] Guiding logic: Based on pre-set rule templates (such as "Please specify the time range" "Need to clarify daily / weekly / monthly aggregation") or large model-generated follow-up questions, guide the user to refine the requirements.
[0168] Example:
[0169] User input: "Statistical sales data and store in database".
[0170] System follow-up: "Please specify the data source type (MySQL / Hive, etc.), statistical dimension (by region / product), and target database name".
[0171] S3: Multi-round interaction optimization.
[0172] In the preferred embodiment, step S3 includes:
[0173] S31, refine questions.
[0174] The system proposes refined questions based on the initial understanding of the task, helping the user to supplement the necessary parameters. For example, the system will ask the user which fields they specifically need to collect, whether they need to filter certain data, and which table in ClickHouse to store in, etc.
[0175] S32, dynamically update requirements.
[0176] As each round of conversation progresses, the user can adjust the requirements, and the system will update the understanding of the task scenario in real time to ensure the accuracy of the requirement description.
[0177] For example, the user's initial requirement: "Collect last week's sales data from MySQL."
[0178] The system asks: "Please specify the specific date range."
[0179] The user corrects: "Change to January 1, 2023 to January 7, 2023."
[0180] The system dynamically updates the time range parameter in the requirement and generates scripts based on the new parameters in subsequent dialogues.
[0181] S33, Real-time Suggestions.
[0182] The system intelligently recommends operation options and processing strategies based on information in the dialogue. For example, in the data cleaning scenario, the system can automatically prompt the user to select "de-duplication" or "fill in missing values" operations.
[0183] S4: Intent Verification and Confirmation.
[0184] In the preferred embodiment, step S4 includes:
[0185] S41, Requirement Confirmation.
[0186] The system will generate a requirement summary containing all the key information provided by the user.
[0187] Example:
[0188] The system generates a requirement summary:
[0189] - Data source: MySQL (address: 192.168.1.100, table: sales) - Collection fields: order_id, amount, date
[0190] - Processing operation: daily aggregation by date field, sum amount
[0191] - Target storage: ClickHouse (table: daily_sales)
[0192] S42, Confirmation Modification.
[0193] If the user has any objections or wishes to modify the requirement description, the system will adjust again until the requirement is fully confirmed.
[0194] Example:
[0195] User feedback: "The aggregation granularity should be weekly, not daily."
[0196] System response:
[0197] Update the "time granularity" in the requirement to "weekly."
[0198] The demand summary is regenerated and the user is required to confirm again.
[0199] S5: configuration script generation.
[0200] In a preferred embodiment, step S5 comprises: the system generates a configuration script. Once the demand is confirmed, the system will automatically generate a configuration script according to the identified task type and details. These configuration scripts will be automatically generated according to different task scenarios, and the format conforms to the specifications of the user's target platform (such as MySQL, ClickHouse, etc.). For example, the system generates a SQL script to collect data from the users table in MySQL, aggregate it by day, and store the result in the daily_users table in ClickHouse, with the aggregation method being summation.
[0201] Based on the same inventive concept, the present application also proposes a big model-based data collection and processing configuration generation system. The implementation of this system can refer to the implementation of the above-mentioned method, and the repeated parts will not be repeated. The term "module" used below can be a combination of software and / or hardware that implements a predetermined function. Although the system described in the following embodiments is preferably implemented in software, hardware, or a combination of software and hardware implementation is also possible and contemplated.
[0202] Figure 2 is a structural diagram of the big model-based data collection and processing configuration generation system of the present application. As shown in Figure 2 , the system comprises:
[0203] The training big model module 101 is used to build in a prompt word template and train a big model; specifically as follows:
[0204] The prompt word template is built in to train the intention recognition ability of the big model;
[0205] Define the mapping relationship between the template key field and the data processing ability, including:
[0206] Define the key field commonly used in data processing tasks;
[0207] Define the data processing ability library, which refers to a group of functional modules or components for processing user demands; each component is associated with a specific task type and key field;
[0208] Map the above-mentioned key field and data processing ability library component.
[0209] The initial demand capture module 102 is used for the user to input natural language demand, and the big model is used for intention recognition and semantic analysis of the input text to preliminarily identify the user's demand, including task type and key parameters.
[0210] The multi-round interaction optimization module 103 is configured to propose a refined question, dynamically update a requirement, and make a real-time suggestion, and to further improve the requirement of the user, obtain more detailed information, and help the system accurately identify the requirement through multi-round interaction with the user.
[0211] The intent verification and confirmation module 104 is configured to, after the multi-round interaction, aggregate all known information, and confirm the final requirement with the user. If the user has any objection to the requirement description or wishes to make a modification, the system will adjust again until the requirement is completely confirmed.
[0212] The configuration script generation module 105 is configured to, once the requirement is confirmed, automatically generate a configuration script in a format conforming to the specification of the target platform of the user according to the identified task type and details.
[0213] It should be noted that although several modules of the large model-based data collection and processing configuration generation system are mentioned in the foregoing detailed description, such division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided into multiple modules.
[0214] Based on the foregoing inventive concept, as shown in Figure 3 The present application also proposes a computer device 200, which comprises a memory 210, a processor 220, and a computer program 230 stored on the memory 210 and executable on the processor 220, wherein the processor 220 implements the foregoing large model-based data collection and processing configuration generation method when executing the computer program 230.
[0215] Based on the foregoing inventive concept, the present application also proposes a computer-readable storage medium, which stores a computer program for executing the foregoing large model-based data collection and processing configuration generation method.
[0216] Due to the advancement of the present application, it can have wide application in the fields of data processing, artificial intelligence application and software development, etc. First of all, in the field of big data processing, the present application can greatly improve the efficiency and accuracy of data processing. Through natural language processing technology, users can simply describe the data processing requirements, and the system can automatically generate execution scripts without manually writing SQL statements or having programming skills. This not only reduces the threshold of data processing, but also reduces the error rate of manual operation and improves efficiency. In addition, the present application can be dynamically adjusted according to user requirements, with high flexibility, which can meet the diversified data processing needs of different users. Secondly, in the field of artificial intelligence application, the present application can be used in intelligent customer service, intelligent assistants and other applications. Through natural language processing technology, the system can have natural language conversation with users, understand the user's intention, and provide intelligent solutions according to the user's needs. This not only improves the user experience, but also improves the work efficiency and reduces the workload of manual customer service. In general, the present application has wide application prospect and market demand, and is expected to play an important role in the fields of big data processing, artificial intelligence application, etc.
[0217] Although the spirit and principles of the present application have been described with reference to several specific embodiments, it should be understood that the present application is not limited to the disclosed specific embodiments, and the division of aspects does not mean that the features in these aspects cannot be combined for benefit, but only for the convenience of expression. The present application is intended to cover various modifications and equivalent arrangements contained in the spirit and scope of the appended claims.
[0218] The scope of protection of the present application should be understood by those skilled in the art that various modifications or changes made on the basis of the technical solutions of the present application without creative labor are still within the scope of protection of the present application.
Claims
1. A large model-based data acquisition and processing configuration generation method, characterized in that, The method comprises: The user inputs a natural language requirement, and a large model is used for intent recognition and semantic analysis on the input text to preliminarily identify the user's requirement, including a task type and a key parameter; The system proposes a refined question, dynamically updates the requirement, and makes a real-time suggestion, and interacts with the user in multiple rounds to further improve the user's requirement and obtain more detailed information; After multiple rounds of interaction, the system will summarize all known information and confirm the final requirement with the user, and if the user has any objection to the requirement description or wishes to make a modification, the system will adjust again until the requirement is completely confirmed; Once the requirement is confirmed, the system will automatically generate a configuration script in a format conforming to the specification of the user's target platform according to the identified task type and details.
2. The method of claim 1, wherein, The method further comprises: built-in prompt word templates are used to train the large model.
3. The method of claim 2, wherein, The built-in prompt word templates are used to train the large model, including: The built-in prompt word templates are used to train the large model's intent recognition capability; A mapping relationship between a template key field and a data processing capability is defined.
4. The method of claim 3, wherein, The mapping relationship between the template key field and the data processing capability includes: Common key fields in data processing tasks are defined; A data processing capability library is defined, which refers to a group of functional modules or components for processing user requirements; each component is associated with a specific task type and a key field; The above key fields are mapped to the data processing capability library components.
5. A large model-based data acquisition and processing configuration generation system, characterized by, The system comprises: An initial requirement capturing module is configured to capture the user's natural language requirement, and a large model is used for intent recognition and semantic analysis on the input text to preliminarily identify the user's requirement, including a task type and a key parameter; A multiple-round interaction optimization module is configured to propose a refined question, dynamically update the requirement, and make a real-time suggestion, and interact with the user in multiple rounds to further improve the user's requirement and obtain more detailed information; An intent verification and confirmation module is configured to, after multiple rounds of interaction, summarize all known information and confirm the final requirement with the user, and if the user has any objection to the requirement description or wishes to make a modification, the system will adjust again until the requirement is completely confirmed; A configuration script generation module is configured to, once the requirement is confirmed, automatically generate a configuration script in a format conforming to the specification of the user's target platform according to the identified task type and details.
6. The large model based data acquisition and processing configuration generation system of claim 5, wherein, The method further comprises: a large model training module is configured to build-in prompt word templates and train the large model.
7. The large model based data acquisition and processing configuration generation system of claim 6, wherein, The large model training module is specifically configured to: The built-in prompt word templates are used to train the large model's intent recognition capability; A mapping relationship between a template key field and a data processing capability is defined.
8. The large model based data acquisition and processing configuration generation system of claim 7, wherein, The mapping relationship between the template key field and the data processing capability includes: Common key fields in data processing tasks are defined; A data processing capability library is defined, which refers to a group of functional modules or components for processing user requirements; each component is associated with a specific task type and a key field; The above key fields are mapped to the data processing capability library components.
9. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method of any one of claims 1-4.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program for executing the method of any one of claims 1-4.