Data modeling interaction method and device and computer readable storage medium
Through natural language processing and intelligent question-and-answer system, the problem of high technical background requirements for traditional data modeling is solved, efficient modeling and accuracy for non-professional personnel is achieved, and technical threshold is lowered.
Patent Information
- Application Number
- CN202510376911.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
The traditional data modeling process has high technical background requirements, which makes it difficult for non-professional personnel to participate, and the modeling efficiency is inefficient, which cannot meet the rapidly changing business needs.
By receiving natural language questions, using natural language processing and intelligent question-and-answer systems, answers and suggestions are generated, and automated processing is reduced through graphical representation and operator configuration, and the technical threshold is reduced and efficiency is improved.
It enables efficient data modeling for non-professional personnel, lowers technical thresholds, improves modeling efficiency and accuracy, and reduces the probability of manual intervention and errors.
Smart Images

Figure CN120336283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and data modeling, and particularly to a data modeling interaction method, device, and computer-readable storage medium. Background Art
[0002] In traditional data modeling scenarios, data modeling work is like a complex maze. It is no easy task for users to successfully navigate through it and complete complex SQL statement configurations and operator information configurations. This process requires users to have a certain degree of technical background knowledge, including a proficient grasp of database principles, SQL syntax, and an in-depth understanding of the functions and applicable scenarios of various operators.
[0003] Taking SQL statement configuration as an example, users not only need to be familiar with the precise usage of various basic query statements, such as SELECT, FROM, WHERE, etc., but also face advanced operations such as complex multi-table joins and subquery nesting. When performing a multi-table join, users need to accurately select the appropriate join type, such as inner join, outer join (left join, right join, full join), etc., according to the logical relationship between the data. A slight mistake may lead to deviations in the data results. Subquery nesting is even more challenging for users' ability to grasp the data hierarchy. How to reasonably nest subqueries in the main query to ensure that the data returned by the subquery can work in cooperation with the main query to obtain accurate results requires a large amount of practical experience and profound technical skills.
[0004] Operator information configuration is equally complex. Different operators have unique functions and applicable scenarios. For example, for a data filtering operator, users need to accurately set filtering conditions to ensure that the filtered data meets business requirements; for an aggregation operator, users need to clarify the selection of aggregation functions and the dimensions of aggregation to achieve effective data statistics and analysis.
[0005] This relatively high technical threshold undoubtedly becomes an insurmountable obstacle for the vast number of potential users. It discourages many people who have a need for data modeling but lack a professional technical background, greatly hindering the popularization of data modeling among a wider population. At the same time, for those users with certain technical capabilities, due to the complex configuration process requiring a large amount of time and effort for debugging and optimization, the modeling efficiency is also greatly reduced, unable to meet the rapidly changing business needs. Summary of the Invention
[0006] In view of this, the present invention provides a data modeling interaction method, device, and computer-readable storage medium, which can improve the efficiency and usability of data modeling, lower the technical threshold, and enable non-professionals to perform data modeling efficiently.
[0007] An embodiment of the present invention provides a data modeling interaction method, and the method includes:
[0008] Receive a natural language question related to data modeling input by the user;
[0009] Utilize natural language processing techniques to parse the natural language question and obtain the user's question intention;
[0010] Based on the knowledge base built into the intelligent question-answering system, which includes common questions about data modeling, best practices, and detailed information on specific data modeling tools and languages, generate accurate answers and suggestions for the user's question intention.
[0011] Further, the method further includes:
[0012] Automatically collect interaction data during the interaction between the user and the system, and at the same time collect the user's active feedback data through the feedback entry;
[0013] Apply data mining and machine learning techniques to extract key features from the collected interaction data and feedback data, analyze the correlation between the user's question patterns, answer satisfaction, and the key features, and obtain the key factors affecting the question-answering effect;
[0014] Based on the key factors affecting the question-answering effect, adopt reinforcement learning or neural network fine-tuning techniques to adjust the intelligent question-answering system model;
[0015] Construct a user portrait based on historical interaction data and feedback data, divide user groups through clustering, formulate personalized question-answering strategies for different groups, or provide personalized assistance according to the user portrait matching strategy.
[0016] Further, the method further includes:
[0017] Parse the data model created by the user, and identify the logical structure and attribute information of the data model;
[0018] Apply a graphics generation algorithm to automatically generate a graphical representation corresponding to the logical structure and attribute information of the data model, and the graphical representation displays the key information of the data model with intuitive graphical elements.
[0019] Further, the method further includes:
[0020] Receive and parse the requirements for various operator configurations input by the user in natural language, and obtain the parsing results of various operator configurations;
[0021] According to the parsing results of various operator configurations, convert the configuration requirements of various operators into corresponding operator configurations;
[0022] Perform error checking and correction on the generated operator configurations.
[0023] Further, the various types of operators include: SQL operators, filtering operators, field adding operators, arithmetic operators, string operators, HTTP operators, and date operators.
[0024] Further, the method further includes:
[0025] Capturing the operation behavior of the user dragging an operator or a data table into the canvas;
[0026] Using machine learning algorithms to extract the functional features and input / output features of the operator, and to extract the structure, data pattern, and semantic features of the data table;
[0027] According to the functional features and input / output features of the operator and the structure, data pattern, and semantic features of the data table, analyzing the potential association relationship between the operator and the data table through an association analysis algorithm;
[0028] According to the predefined optimization objective, determining the best association relationship from the potential association relationships, and presenting suggestions to the user in an intuitive manner.
[0029] Further, the method further includes:
[0030] Capturing the operation behavior of the user dragging multiple data tables into the canvas;
[0031] Applying data fusion technology to analyze the structure information, data statistical features, and semantic information of each data table, and identifying the association relationship between each data table through a pattern matching algorithm;
[0032] According to the association relationship between each data table, combining the objectives and strategies of data fusion, generating a merging suggestion for each data table, and visually presenting the merging suggestion to the user.
[0033] Further, the method further includes:
[0034] Real-time monitoring of operator addition operations;
[0035] When a new operator is detected to be added, analyzing the field information of the data table corresponding to the new operator and the output data features of the upstream operator of the new operator;
[0036] According to the analysis results, combining with a preset rule library, identifying and automatically adding the configuration items of the new operator.
[0037] In a second aspect, an embodiment of the present invention provides a data modeling interaction device, and the device includes:
[0038] A receiving module, configured to receive a natural language question related to data modeling input by a user;
[0039] A parsing module, configured to parse the natural language problem by using natural language processing technology to obtain the user's question intention;
[0040] A generation module, configured to generate accurate answers and suggestions for the user's question intention according to a knowledge base built into the intelligent question-answering system, where the knowledge base includes common problems in data modeling, best practices, and detailed information about specific data modeling tools and languages.
[0041] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the method described in any item of the first aspect when running.
[0042] In a fourth aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, characterized in that a computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any item of the first aspect.
[0043] The technical solution provided by the present invention receives a natural language problem related to data modeling input by a user, parses the natural language problem by using natural language processing technology to obtain the user's question intention, and generates accurate answers and suggestions for the user's question intention according to a knowledge base built into the intelligent question-answering system. On the one hand, it can improve the usability of data modeling, lower the professional threshold, and non-professional personnel can operate through natural language interaction; simplify the process, automatically associate, combine tables, and add configuration items, reducing manual operations. On the second hand, it can enhance the intelligence and automation of data modeling. The intelligent question-answering system uses natural language processing to understand the user's intention, provides answers and suggestions in combination with the knowledge base, and optimizes the question-answering ability through learning to provide personalized help, realizing efficient intelligent interaction, configuring various operators in natural language, and the system automatically converts the configuration; automatically associates, combines tables, and adds configuration items to automatically execute data modeling tasks, improving the processing efficiency, reducing manual intervention, and lowering the error probability; in the third aspect, the error detection and correction function of configuring operators in natural language ensures that the generated SQL statements and other operator configurations are accurate and effective, guaranteeing the accuracy and efficiency of data processing. At the same time, automatically adding configuration items analyzes according to data tables and the output of upstream operators to ensure the smooth and correct data flow and avoid data processing errors caused by improper configuration. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 is a flowchart of a data modeling interaction method provided by Embodiment 1 of the present invention.
[0045] Figure 2 is a schematic structural diagram of a data modeling interaction device provided by Embodiment 8 of the present invention.
[0046] Figure 3 It is a schematic structural diagram of an electronic device provided in the ninth embodiment of the present invention. Detailed implementation manners
[0047] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0048] Embodiment 1
[0049] See Figure 1 , Figure 1 It is a flowchart of a data modeling interaction method provided in Embodiment 1 of the present invention. The data modeling interaction method includes the following steps:
[0050] Step 101: Receive a natural language question related to data modeling input by a user.
[0051] In the present invention, a dedicated input area is set on the front-end interface of the intelligent question-answering system. This area is usually presented in the form of a text box, where the user can input any natural language question related to data modeling. For example, the user may input questions such as "How to select a suitable algorithm for predicting sales data in data modeling?" or "What are the commonly used libraries for data modeling using Python?" etc. The intelligent question-answering system obtains the content input by the user in real time by listening to the input event of the user in this input box. When the user clicks the "Submit" button or presses the Enter key, the system sends the natural language question input by the user to the backend for processing.
[0052] In this step, to ensure accurate reception of various forms of user input, the input area needs to have good compatibility and support access from different devices (such as computers, tablets, mobile phones) and browsers. At the same time, considering the convenience of user input, some auxiliary functions, such as auto-completion or input prompts, can be added to help the user express the question more accurately.
[0053] Step 102: Use natural language processing technology to parse the natural language question and obtain the user's question intention.
[0054] In this step, the intelligent Q&A system first uses a word segmentation tool to split the natural language question input by the user into individual words or phrases. For example, for the question "How to select a suitable algorithm for predicting sales data in data modeling?", it may be split into words such as "How", "in", "data modeling", "select", "suitable", "algorithm", "predict", "sales data", etc. Common word segmentation algorithms include rule-based word segmentation, statistics-based word segmentation, and deep learning word segmentation methods. The system will select a suitable word segmentation tool according to the actual situation.
[0055] After that, each segmented word is tagged with a part of speech to determine its grammatical role in the sentence, such as noun, verb, adjective, etc. For example, "data modeling" may be tagged as a noun, and "select" as a verb. This helps to understand the sentence structure and semantic relationships.
[0056] After that, by analyzing the grammatical relationships between words, a syntactic structure tree of the sentence is constructed. For example, it is determined that "select" is the predicate verb, "algorithm" is the object, and "How" is an interrogative adverb modifying the action of "select", so as to clearly understand the grammatical level of the sentence.
[0057] Finally, combined with the semantic knowledge in the knowledge base and the pre-trained language model, the semantics of the sentence are understood. For example, it is recognized that the core of the user's question is about algorithm selection in the data modeling scenario and involves the application scenario of predicting sales data. The system will use word vector models (such as Word2Vec, GloVe) to map words into a vector space and calculate the similarity between vectors to judge the semantic relevance of words and sentences, so as to more accurately grasp the user's question intention.
[0058] Natural language is diverse and flexible, and the same meaning may have multiple expressions. Therefore, a large amount of corpus data is needed to train the natural language processing model to improve the model's understanding ability of various expressions. In addition, the training and inference efficiency of the model are also very important. It is necessary to minimize the processing time under the premise of ensuring accuracy.
[0059] Step 103: Based on the knowledge base built into the intelligent Q&A system, which contains common questions, best practices in data modeling, and detailed information on specific data modeling tools and languages, generate accurate answers and suggestions for the user's question intention.
[0060] In this step, the user question intention parsed is matched with the common questions in the knowledge base. The questions in the knowledge base are usually stored in a normalized form, and the system will find the knowledge base question that best matches the user's question by calculating similarities and other methods. For example, for the user's question "How to select a suitable algorithm in data modeling to predict sales data?", the system may find a similar question in the knowledge base such as "In data modeling for sales data prediction, how to select a suitable algorithm?"
[0061] Once a matching question is found, obtain the corresponding answers and suggestions from the knowledge base. The answers may include detailed step-by-step instructions, relevant case analyses, recommended tools or algorithms, etc. For example, the answer may mention that "When predicting sales data, a suitable algorithm can be selected according to the characteristics of the data. If the data shows a linear trend, the linear regression algorithm may be a good choice; if the data has complex non-linear relationships, algorithms such as decision trees, random forests, or neural networks may be more suitable. At the same time, the characteristics of seasonality and periodicity of historical sales data can be referred to for algorithm adjustment."
[0062] Preferably, if there is no exactly matching question in the knowledge base, the system may utilize relevant knowledge fragments in the knowledge base and combine natural language generation technology to optimize and combine the answers to generate answers that better meet the user's needs. For example, extract relevant information from different best practice cases and organize them into a complete and targeted answer. In addition, the system may also make personalized adjustments to the answers according to the user's historical interaction records and preferences. For example, if the user has often focused on Python-related content before, the answer will mention more methods and libraries for data modeling using Python.
[0063] Embodiment 2
[0064] A data modeling interaction method provided by Embodiment 2 of the present invention, the data modeling interaction method includes the following steps:
[0065] Step 201, receive a natural language question related to data modeling input by the user.
[0066] Step 202, use natural language processing technology to parse the natural language question to obtain the user question intention.
[0067] Step 203, according to the knowledge base built into the intelligent question and answer system, the knowledge base contains common questions about data modeling, best practices, and detailed information on specific data modeling tools and languages, and generate accurate answers and suggestions for the user question intention.
[0068] The above steps 201 - step 203 can be understood with reference to Embodiment 1 above, and will not be elaborated in this embodiment.
[0069] Step 204: Automatically collect interaction data when the user interacts with the system, and at the same time collect the user's active feedback data through the feedback entry.
[0070] During the process of the user interacting with the intelligent Q&A system, the system will automatically collect various types of interaction data, including the questions entered by the user, the answers given by the system, the stay time of the user after viewing the answers, whether to click on further relevant links, and other information. These data can reflect the detailed behavior of the user's interaction with the system.
[0071] At the same time, the system sets up a feedback entry to encourage users to actively provide feedback data. For example, users can express their satisfaction with the answers and whether they have improvement suggestions by filling out questionnaires, submitting evaluations, etc. These two data collection methods complement each other and comprehensively record the interaction between the user and the system.
[0072] Step 205: Use data mining and machine learning techniques to extract key features from the collected interaction data and feedback data, analyze the association between the user's question pattern, answer satisfaction and the key features, and obtain the key factors affecting the Q&A effect.
[0073] Use data mining techniques to screen out valuable information from the large amount of collected interaction data and feedback data, and convert it into quantifiable features. For example, extract features such as question type and complexity from the user's questions, and extract satisfaction scores and feedback opinion keywords from the feedback data.
[0074] With the help of machine learning algorithms, analyze the internal relationship between these key features and the user's question pattern (such as question frequency, common question types) and answer satisfaction. For example, through correlation analysis, find out which question features are closely related to low-satisfaction answers, so as to determine the key factors affecting the Q&A effect, such as the lack of knowledge in a specific field, the answer expression method, etc.
[0075] Step 206: Based on the key factors affecting the Q&A effect, use reinforcement learning or neural network fine-tuning technology to adjust the intelligent Q&A system model.
[0076] Based on the key factors affecting the Q&A effect determined in Step 205, use reinforcement learning technology. Regard the Q&A system as an agent, and through continuous interaction with the user to obtain rewards (such as positive rewards for high user satisfaction) or punishments (such as negative rewards for user dissatisfaction), let the system learn how to give better answers and gradually optimize the model.
[0077] Alternatively, neural network fine-tuning technology can be used to fine-tune the model parameters based on the pre-trained neural network model according to key factors. For example, if it is found that the answers to specific types of questions are inaccurate, the parameters of the model part for handling such questions are adjusted so that the model can better handle these situations and improve the accuracy and quality of the Q&A.
[0078] Step 207: Construct a user profile based on historical interaction data and feedback data, divide user groups through clustering, formulate personalized Q&A strategies for different groups, or provide personalized assistance according to the user profile matching strategy.
[0079] Construct a user profile based on historical interaction data and feedback data. This profile contains various information such as the user's interest areas, knowledge level, usage habits, etc. For example, for a user who frequently asks questions related to algorithms in data modeling, their profile may highlight the interest point in algorithms.
[0080] Divide users into different groups through clustering algorithms, such as beginner groups, professional groups, etc. According to the characteristics of different groups, formulate personalized Q&A strategies. For example, provide more detailed and basic explanations and examples for beginners, and provide more in-depth and cutting-edge content for professionals.
[0081] In addition, specific strategies can also be matched according to a single user profile. For example, according to the data modeling tool preferred by the user, relevant solutions of the tool are preferentially recommended in the answer to provide personalized assistance to the user and improve the user experience.
[0082] Embodiment III
[0083] Embodiment III of the present invention provides a data modeling interaction method, and the data modeling interaction method includes the following steps:
[0084] Step 301: Receive a natural language question related to data modeling input by the user.
[0085] Step 302: Use natural language processing technology to parse the natural language question to obtain the user's question intention.
[0086] Step 303: Based on the knowledge base built into the intelligent Q&A system, the knowledge base contains common questions in data modeling, best practices, and detailed information on specific data modeling tools and languages. For the user's question intention, generate accurate answers and suggestions.
[0087] Step 304: Automatically collect interaction data during the interaction between the user and the system, and at the same time collect user initiative feedback data through the feedback entry.
[0088] Step 305: Use data mining and machine learning techniques to extract key features from the collected interaction data and feedback data, analyze the association between the user's question patterns, answer satisfaction, and the key features, and obtain the key factors affecting the Q&A effect.
[0089] Step 306: Based on the key factors affecting the Q&A effect, use reinforcement learning or neural network fine-tuning techniques to adjust the intelligent Q&A system model.
[0090] Step 307: Construct a user profile based on historical interaction data and feedback data, divide user groups through clustering, formulate personalized Q&A strategies for different groups, or provide personalized assistance according to the user profile matching strategy.
[0091] The above Steps 301 - 307 can be understood with reference to Embodiment 2 and will not be elaborated here.
[0092] Step 308: Analyze the data model created by the user and identify the logical structure and attribute information of the data model.
[0093] This step aims to deeply understand the data model created by the user. The system will analyze the data model, just like disassembling a complex machine to figure out the structure and characteristics of each part. Here, the "logical structure" refers to the organizational relationship between data elements in the data model, such as a hierarchical structure, a network structure, or other types. For example, in an e-commerce data model, the association method between products, orders, and users constitutes the logical structure. And the "attribute information" refers to the characteristics of each data element. For example, the data elements of a product may have attributes such as name, price, and inventory. Through analysis, the system can accurately grasp the internal composition of the data model and prepare for the subsequent generation of a graphical representation.
[0094] Step 309: Use a graphics generation algorithm to automatically generate a graphical representation corresponding to the logical structure and attribute information of the data model, and the graphical representation displays the key information of the data model with intuitive graphical elements.
[0095] After obtaining the logical structure and attribute information of the data model, the system uses a specific graphics generation algorithm to create a drawing. This algorithm is like a highly skilled painter who creates an intuitive graph based on the information of the data model. It will display the key information in the data model with various graphical elements. For example, it may use rectangles to represent different data elements, arrows to represent the relationships between them, and different colors or annotations to reflect the attribute information. The finally generated graph allows users to clearly see the core content of the data model at a glance, greatly reducing the difficulty of understanding the data model, just like quickly understanding the layout of a complex city through a map.
[0096] Embodiment 4
[0097] Embodiment 4 of the present invention provides a data modeling interaction method, and the data modeling interaction method includes the following steps:
[0098] Step 401, receive a natural language question related to data modeling input by a user.
[0099] Step 402, use natural language processing technology to parse the natural language question to obtain the user's question intention.
[0100] Step 403, according to the knowledge base built into the intelligent question-answering system, where the knowledge base includes common problems in data modeling, best practices, and detailed information about specific data modeling tools and languages, generate accurate answers and suggestions for the user's question intention.
[0101] Step 404, automatically collect interaction data during the interaction between the user and the system, and at the same time collect user initiative feedback data through the feedback entry.
[0102] Step 405, use data mining and machine learning technologies to extract key features from the collected interaction data and feedback data, analyze the association between the user's question pattern, answer satisfaction, and the key features, and obtain the key factors affecting the question-answering effect.
[0103] Step 406, based on the key factors affecting the question-answering effect, use reinforcement learning or neural network fine-tuning technology to adjust the intelligent question-answering system model.
[0104] Step 407, construct a user portrait according to historical interaction data and feedback data, divide user groups through clustering, formulate personalized question-answering strategies for different groups, or provide personalized help according to the strategy matching of the user portrait.
[0105] Step 408, parse the data model created by the user, and identify the logical structure and attribute information of the data model.
[0106] Step 409, use a graphics generation algorithm to automatically generate a graphical representation corresponding to the logical structure and attribute information of the data model, and the graphical representation displays the key information of the data model with intuitive graphical elements.
[0107] The above steps 401-step 409 can be understood with reference to Embodiment 3, and will not be elaborated here.
[0108] Step 410, receive and parse the requirements for various operator configurations input by the user in natural language, and obtain the parsing results of various operator configurations.
[0109] The system first waits for the user to describe various operator configuration requirements in natural language. For example, the user may input "Filter out employee data where the age is greater than 30 years old", which involves the configuration requirements of the filtering operator. After receiving it, the system uses natural language processing techniques, such as word segmentation, part-of-speech tagging, syntactic analysis, and semantic understanding, to convert the natural language into information that the computer can understand. For example, it recognizes that "age greater than 30 years old" is the filtering condition and "employee data" is the object of operation, so as to obtain the parsing results of various operator configurations and clarify what operations the user actually wants to perform on the data.
[0110] In this embodiment, various operators may include: SQL operator, filtering operator, adding field operator, arithmetic operator, string operator, HTTP operator, and date operator.
[0111] Step 411: According to the parsing results of the configurations of various operators, convert the configuration requirements of various operators into corresponding operator configurations.
[0112] Based on the parsing results obtained in the previous step, the system converts the operator configuration requirements described in natural language into actual executable operator configurations according to preset rules and mapping relationships. For example, the requirement "Filter out employee data where the age is greater than 30 years old" is converted into the specific parameter settings of the filtering operator, that is, specifying the filtering field as "age", the operator as "greater than", the value as "30", and determining the data table for the operation as "employee data table", completing the conversion from the user's intention to the actual configuration.
[0113] Step 412: Perform error checking and correction on the generated operator configuration.
[0114] After generating the operator configuration, the system will perform a comprehensive check on it. Through built-in verification mechanisms such as syntax rules, logical rules, and data type constraints, it checks whether there are errors in the configuration. For example, it checks whether the filtering field actually exists in the corresponding data table and whether the operator matches the data type. If an error is found, the system will try to perform automatic correction according to the knowledge base or preset correction strategies. For example, if it is found that the spelling of the specified filtering field name is incorrect, the system will search for similar fields in the knowledge base and correct them to ensure that the operator configuration can be executed accurately.
[0115] Embodiment Five
[0116] Embodiment Five of the present invention provides a data modeling interaction method, and the data modeling interaction method includes the following steps:
[0117] Step 501: Receive a natural language question related to data modeling input by the user.
[0118] Step 502: Parse the natural language question using natural language processing technology to obtain the user's question intention.
[0119] Step 503: Based on the knowledge base built into the intelligent question - answering system, which includes common problems in data modeling, best practices, and detailed information about specific data modeling tools and languages, generate accurate answers and suggestions for the user's question intention.
[0120] Step 504: Automatically collect interaction data when the user interacts with the system, and at the same time collect the user's active feedback data through the feedback entry.
[0121] Step 505: Use data mining and machine learning technologies to extract key features from the collected interaction data and feedback data, analyze the correlation between the user's question pattern, answer satisfaction, and the key features, and obtain the key factors affecting the question - answering effect.
[0122] Step 506: Based on the key factors affecting the question - answering effect, use reinforcement learning or neural network fine - tuning technology to adjust the intelligent question - answering system model.
[0123] Step 507: Construct a user profile based on historical interaction data and feedback data, divide user groups through clustering, formulate personalized question - answering strategies for different groups, or provide personalized help according to the user - profile matching strategy.
[0124] Step 508: Parse the data model created by the user to identify the logical structure and attribute information of the data model.
[0125] Step 509: Use a graphics generation algorithm to automatically generate a graphical representation corresponding to the logical structure and attribute information of the data model, and the graphical representation displays the key information of the data model with intuitive graphical elements.
[0126] Step 510: Receive and parse the user's requirements for various operator configurations input in natural language to obtain the parsing results of various operator configurations.
[0127] Step 511: According to the parsing results of various operator configurations, convert the configuration requirements of various operators into corresponding operator configurations.
[0128] Step 512: Perform error checking and correction on the generated operator configurations.
[0129] The above steps 501 - step 512 can be understood by referring to steps 401 - step 412 in Embodiment 4 above respectively, and will not be elaborated here.
[0130] Step 513: Capture the user's operation behavior of dragging an operator or a data table into the canvas.
[0131] The system monitors the user's operations on the interface at all times. When the user uses the mouse to drag and drop an operator (such as a data filtering operator, an aggregation operator, etc.) or a data table from a specific area to the canvas area, the system can immediately detect and record this operation behavior, providing trigger conditions for subsequent analysis.
[0132] Step 514: Utilize a machine learning algorithm to extract the functional features and input and output features of the operator, and extract the structure, data model, and semantic features of the data table.
[0133] For the dragged-in operators, the machine learning algorithm is used to analyze the internal code logic, function description documents and other information to extract the functional features of the operators. For example, the function of the data filtering operator is to filter data according to the set conditions. At the same time, the input and output features of the operator are clarified, that is, what kind of data it needs as input and what form of data it will output after processing.
[0134] For the dragged-in data table, the machine learning algorithm is also used to obtain its structural features from the metadata of the data table, such as which fields it contains and the data types of the fields. The data pattern features are extracted through statistical analysis of the data in the table, such as whether the data has a certain distribution pattern. In addition, semantic features are extracted in combination with domain knowledge to understand the business meaning represented by the data table. For example, the semantics of a data table named "sales records" is related to the sales business.
[0135] Step 515: Analyze the potential association relationship between the operator and the data table through an association analysis algorithm according to the functional characteristics and input and output characteristics of the operator and the structure, data model and semantic characteristics of the data table.
[0136] Based on the various features of the operators and data tables extracted in the previous step, the association analysis algorithm is used for in-depth mining. The algorithm looks for potential connections between operator functions and data table structures, data patterns, and semantics. For example, an aggregation operator used to calculate average values may have a potential connection with a data table containing numeric fields, because numeric fields are suitable for average value calculation. Through this analysis, all possible associations between operators and data tables are found.
[0137] Step 516: According to the predefined optimization goal, determine the best association relationship from the potential association relationships, and make suggestions to the user in an intuitive manner.
[0138] The system has pre-set some optimization goals, such as maximizing data processing efficiency, ensuring data accuracy, etc. According to these goals, the potential association relationships obtained in step 515 are evaluated and screened. The best association relationship that best meets the optimization goals is determined from among numerous potential relationships. For example, an association method that can make the data processing process the simplest and the results the most accurate is selected. Then, this best association relationship is presented and recommended to the user in an intuitive way such as through interface prompts and visual line connections, facilitating the user's subsequent operations.
[0139] Embodiment Six
[0140] Embodiment Six of the present invention provides a data modeling interaction method, and the data modeling interaction method includes the following steps:
[0141] Step 601, receive a natural language question related to data modeling input by the user.
[0142] Step 602, use natural language processing technology to parse the natural language question to obtain the user's question intention.
[0143] Step 603, based on the knowledge base built into the intelligent question-answering system, where the knowledge base contains common questions about data modeling, best practices, and detailed information about specific data modeling tools and languages, generate accurate answers and suggestions for the user's question intention.
[0144] Step 604, automatically collect interaction data during the interaction between the user and the system, and at the same time collect user-initiated feedback data through the feedback entry.
[0145] Step 605, use data mining and machine learning technologies to extract key features from the collected interaction data and feedback data, analyze the association between the user's question pattern, answer satisfaction, and the key features, and obtain the key factors affecting the question-answering effect.
[0146] Step 606, based on the key factors affecting the question-answering effect, use reinforcement learning or neural network fine-tuning technology to adjust the intelligent question-answering system model.
[0147] Step 607, construct a user profile based on historical interaction data and feedback data, divide user groups through clustering, formulate personalized question-answering strategies for different groups, or provide personalized help according to the user profile matching strategy.
[0148] Step 608, parse the data model created by the user, and identify the logical structure and attribute information of the data model.
[0149] Step 609: Automatically generate a graphical representation corresponding to the logical structure and attribute information of the data model using a graphics generation algorithm. The graphical representation displays the key information of the data model with intuitive graphical elements.
[0150] Step 610: Receive and parse the requirements for various operator configurations input by the user in natural language to obtain the parsing results of various operator configurations.
[0151] Step 611: Convert the configuration requirements of various operators into corresponding operator configurations according to the parsing results of various operator configurations.
[0152] Step 612: Perform error checking and correction on the generated operator configurations.
[0153] Step 613: Capture the operation behavior of the user dragging an operator or a data table into the canvas.
[0154] Step 614: Use machine learning algorithms to extract the functional features and input / output features of the operator, and extract the structure, data pattern, and semantic features of the data table.
[0155] Step 615: Analyze the potential association relationships between the operator and the data table through an association analysis algorithm based on the functional features and input / output features of the operator and the structure, data pattern, and semantic features of the data table.
[0156] Step 616: Determine the best association relationship from the potential association relationships according to the predefined optimization goal and provide suggestions to the user in an intuitive manner.
[0157] The above steps 601 - 616 can be understood by referring to steps 501 - 516 in the fifth embodiment above, and will not be elaborated here.
[0158] Step 617: Capture the operation behavior of the user dragging multiple data tables into the canvas.
[0159] The system continuously monitors the actions of the user on the operation interface. When the user drags two or more data tables from the resource area to the canvas area through an input device such as a mouse, the system will quickly perceive and record this operation behavior.
[0160] Step 618: Use data fusion technology to analyze the structure information, data statistical features, and semantic information of each data table, and identify the association relationships between each data table through a pattern matching algorithm.
[0161] Using data fusion technology, the system deeply analyzes various aspects of information in each data table. At the structural level, it understands what fields the data table contains, the data types of the fields, and the arrangement order, etc.; it statistically analyzes the characteristics of the data, such as the mean, maximum and minimum values, and distribution of numerical fields; it mines semantic information to clarify the meaning of the data table in the business scenario. For example, the "customer information table" mainly records customer-related materials.
[0162] With the help of pattern matching algorithms, the system searches for similar or relevant patterns in the information of these data tables to identify the association relationships between them. For example, two tables may both contain the "customer ID" field. By analyzing this field, the system can determine that these two tables may be associated at the customer dimension. For instance, one table records customer basic information, and the other table records customer order information, and they are linked through the "customer ID".
[0163] Step 619: According to the association relationships between the data tables, combined with the goals and strategies of data fusion, generate merge suggestions for each data table, and visually display the merge suggestions to the user.
[0164] Based on the identified association relationships between the data tables, and at the same time combined with the goal of data fusion (such as integrating data for more comprehensive analysis) and the preset strategy (such as merging the same fields according to a certain rule), the system generates specific suggestions on how to merge these data tables. For example, it is recommended to use the "customer ID" as the connection key to merge the "customer information table" and the "customer order information table" to obtain the complete information of customers and their orders.
[0165] To enable the user to more intuitively understand the merge suggestions, the system presents them in a visual way. It may use a graphical interface to connect relevant data tables with lines and mark the associated fields and merge methods; or display the comparison of the data structures before and after the merge in tabular form to help the user clearly see the effect after the merge, so as to facilitate the user to make a decision on whether to adopt the merge suggestion.
[0166] Embodiment Seven
[0167] Embodiment Seven of the present invention provides a data modeling interaction method, and the data modeling interaction method includes the following steps:
[0168] Step 701: Receive a natural language question related to data modeling input by the user.
[0169] Step 702: Use natural language processing technology to parse the natural language question to obtain the user's question intention.
[0170] Step 703: Based on the knowledge base built into the intelligent Q&A system, which contains common problems in data modeling, best practices, and detailed information on specific data modeling tools and languages, generate accurate answers and suggestions for the user's question intention.
[0171] Step 704: Automatically collect interaction data during the interaction between the user and the system, and at the same time collect the user's active feedback data through the feedback entry.
[0172] Step 705: Use data mining and machine learning techniques to extract key features from the collected interaction data and feedback data, analyze the correlation between the user's question pattern, answer satisfaction, and the key features, and obtain the key factors affecting the Q&A effect.
[0173] Step 706: Based on the key factors affecting the Q&A effect, use reinforcement learning or neural network fine-tuning techniques to adjust the intelligent Q&A system model.
[0174] Step 707: Build a user profile based on historical interaction data and feedback data, divide the user groups through clustering, formulate personalized Q&A strategies for different groups, or provide personalized help according to the user profile matching strategy.
[0175] Step 708: Parse the data model created by the user, and identify the logical structure and attribute information of the data model.
[0176] Step 709: Use a graphics generation algorithm to automatically generate a graphical representation corresponding to the logical structure and attribute information of the data model, and the graphical representation displays the key information of the data model with intuitive graphical elements.
[0177] Step 710: Receive and parse the user's requirements for various operator configurations input in natural language, and obtain the parsing results of various operator configurations.
[0178] Step 711: According to the parsing results of various operator configurations, convert the configuration requirements of various operators into corresponding operator configurations.
[0179] Step 712: Perform error checking and correction on the generated operator configurations.
[0180] Step 713: Capture the user's operation behavior of dragging an operator or a data table into the canvas.
[0181] Step 714: Use machine learning algorithms to extract the functional features and input / output features of the operator, and extract the structure, data pattern, and semantic features of the data table.
[0182] Step 715: Analyze the potential association relationship between the operator and the data table through an association analysis algorithm based on the functional characteristics, input and output characteristics of the operator, and the structure, data pattern, and semantic characteristics of the data table.
[0183] Step 716: Determine the best association relationship from the potential association relationships according to the predefined optimization goal, and give suggestions to the user in an intuitive way.
[0184] Step 717: Capture the operation behavior of the user dragging multiple data tables into the canvas.
[0185] Step 718: Use data fusion technology to analyze the structure information, data statistical characteristics, and semantic information of each data table, and identify the association relationships between each data table through a pattern matching algorithm.
[0186] Step 719: Generate merge suggestions for each data table according to the association relationships between each data table, combined with the goals and strategies of data fusion, and display the merge suggestions to the user in a visual way.
[0187] The above Steps 701 - Step 719 can be understood with reference to Steps 601 - Step 619 in Embodiment 6 above respectively, and will not be elaborated here.
[0188] Step 720: Monitor the operator addition operation in real time. When a new operator is detected to be added, analyze the field information of the data table corresponding to the new operator and the output data characteristics of the upstream operator of the new operator.
[0189] The system is always in a "listening" state, continuously paying attention to the operator addition operation. Whether the user manually adds a new operator or the system automatically introduces a new operator, once a new operator is added to the data processing flow, the system will immediately detect it.
[0190] The system then analyzes two types of key information related to the new operator. On the one hand, for the data table that the new operator is about to process, carefully analyze its field information, including details such as the name of each field, data type (such as integer, string, date, etc.), field length, and whether null values are allowed. On the other hand, analyze the characteristics of the output data of the upstream operator of the new operator, such as the data format (whether it is a table, an array, or other forms), value range (for numerical data, its maximum and minimum values), data distribution (such as whether it conforms to a normal distribution), etc. This information is crucial for understanding the input conditions and data processing requirements of the new operator.
[0191] Step 721: Identify and automatically add the configuration items of the new operator according to the analysis results, combined with the preset rule library.
[0192] Based on the analysis results of the data table field information and the output data characteristics of the upstream operator in the previous step, the system also refers to a preset rule library. The rule library stores various configuration rules required by operators under different data conditions, such as data type conversion rules, default value setting rules, field mapping rules, etc. By matching the analysis results with the rules in the rule library, the system identifies the configuration items necessary to ensure the correct operation of the new operator and the smooth flow of the data stream. For example, if the analysis finds that the date format output by the upstream operator is inconsistent with the format expected by the new operator, and there is a corresponding date format conversion rule in the rule library, the system identifies the configuration item that needs to add the date format conversion accordingly.
[0193] After identifying the configuration items, the system automatically adds these configuration items to the new operator. This process does not require manual operation by the user, and the system will directly write the corresponding settings into the configuration parameters of the new operator. For example, in the parameter setting area of the operator, the specific method of data type conversion is automatically filled in, and appropriate default values are set for fields that may be empty, so that the new operator can process data accurately and ensure the continuity and correctness of the entire data processing process.
[0194] Embodiment VIII
[0195] See Figure 2 , Figure 2 FIG. is a schematic structural diagram of a data modeling interaction device provided in Embodiment VIII of the present invention. The device includes:
[0196] A receiving module 81, configured to receive a natural language question related to data modeling input by a user.
[0197] An analysis module 82, configured to use natural language processing technology to analyze the natural language question and obtain the user's question intention.
[0198] A generating module 83, configured to generate accurate answers and suggestions according to the knowledge base built into the intelligent question answering system. The knowledge base includes common problems in data modeling, best practices, and detailed information about specific data modeling tools and languages, for the user's question intention.
[0199] It should be noted that the data modeling interaction device in the embodiments of the present invention and the data modeling interaction method in the above embodiments belong to the same inventive concept. Technical details not described in detail in this device can be referred to the relevant descriptions of the method above and will not be elaborated here.
[0200] In addition, an embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is set to execute the method described above when running.
[0201] Figure 3The structural schematic diagram of an electronic device 10 that can be used to implement the embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0202] As Figure 3 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0203] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0204] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the idle detection method.
[0205] In some embodiments, the idle detection method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the idle detection method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the idle detection method by any other suitable means (e.g., by means of firmware).
[0206] The various implementations of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0207] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0208] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0209] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0210] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0211] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of large management difficulty and weak business scalability existing in traditional physical hosts and VPS services.
[0212] It should be understood that various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0213] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data modeling interaction method, characterized in that, The method includes: Receiving a natural language question related to data modeling input by a user; Using natural language processing technology to parse the natural language question to obtain the user's question intention; According to the knowledge base built into the intelligent question answering system, the knowledge base contains common questions about data modeling, best practices, and detailed information about specific data modeling tools and languages, and generating accurate answers and suggestions for the user's question intention.
2. The method according to claim 1, wherein The method further includes: Automatically collecting interaction data when the user interacts with the system, and at the same time collecting user-initiated feedback data through a feedback entry; Using data mining and machine learning technologies to extract key features from the collected interaction data and feedback data, analyzing the correlation between the user's question pattern, answer satisfaction, and the key features, and obtaining the key factors affecting the question answering effect; Based on the key factors affecting the question answering effect, using reinforcement learning or neural network fine-tuning technology to adjust the intelligent question answering system model; Constructing a user profile based on historical interaction data and feedback data, dividing user groups through clustering, formulating personalized question answering strategies for different groups, or providing personalized help according to the user profile matching strategy.
3. The method according to claim 1, wherein The method further includes: Parsing the data model created by the user, and identifying the logical structure and attribute information of the data model; Using a graphic generation algorithm to automatically generate a graphical representation corresponding to the logical structure and attribute information of the data model, and the graphical representation displays the key information of the data model with intuitive graphic elements.
4. The method according to claim 1, characterized in that, The method further includes: Receiving and parsing the requirements for various operator configurations input by the user in natural language to obtain the parsing results of various operator configurations; According to the parsing results of various operator configurations, converting the configuration requirements of various operators into corresponding operator configurations; Performing error checking and correction on the generated operator configurations.
5. The method according to claim 4, wherein The various operators include: SQL operator, filtering operator, adding field operator, arithmetic operator, string operator, HTTP operator, and date operator.
6. The method according to claim 1, characterized in that, The method further includes: Capturing the operation behavior of the user dragging an operator or a data table into the canvas; Using a machine learning algorithm to extract the functional features and input / output features of the operator, and extracting the structure, data pattern, and semantic features of the data table; According to the functional features and input / output features of the operator and the structure, data pattern, and semantic features of the data table, analyzing the potential association relationship between the operator and the data table through an association analysis algorithm; According to the predefined optimization goal, determining the best association relationship from the potential association relationships, and presenting suggestions to the user in an intuitive manner.
7. The method according to claim 1, wherein The method further includes: Capturing the operation behavior of the user dragging multiple data tables into the canvas; Using data fusion technology to analyze the structural information, data statistical features, and semantic information of each data table, and identifying the association relationship between each data table through a pattern matching algorithm; According to the association relationship between each data table, combining the goal and strategy of data fusion, generating a merger suggestion for each data table, and visually displaying the merger suggestion to the user.
8. The method according to claim 1, wherein The method further includes: Real-time monitor the operator addition operation. When a new operator is detected to be added, analyze the field information of the data table corresponding to the new operator and the output data characteristics of the upstream operator of the new operator; According to the analysis results, in combination with a preset rule library, identify and automatically add the configuration items of the new operator.
9. A data modeling interaction device, characterized in that, The device includes: A receiving module, configured to receive a natural language question related to data modeling input by a user; An analysis module, configured to analyze the natural language question by using natural language processing technology to obtain the user's question intention; A generating module, configured to generate accurate answers and suggestions for the user's question intention based on the knowledge base built in the intelligent question and answer system, where the knowledge base includes common problems in data modeling, best practices, and detailed information about specific data modeling tools and languages.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium, wherein the computer program is set to execute the method according to any one of claims 1 to 8 when running.