A log modeling method based on semantic recommendation

Through the log modeling method based on semantic recommendation, deep learning and semantic inference technology are used to deeply analyze and model log data, solving the problem of insufficient generalization ability and adaptability of log modeling in the existing technology, and achieving efficient and accurate log data processing and intuitive user interaction experience.

CN119537170BActive Publication Date: 2025-07-01COLASOFT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411579676.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-07
Publication Date
2025-07-01
Estimated Expiration
2044-11-07

AI Technical Summary

Technical Problem

When existing log modeling technologies process complex log data in vertical fields, they lack generalization capabilities and adaptability, resulting in limited model accuracy, requiring a large amount of training data and computing resources, which is costly, and at the same time, user experience and popularity are affected.

Method used

The log modeling method based on semantic recommendation is adopted, and deep learning technology combined with semantic inference algorithms is used to deeply analyze the log data, automatically identify and understand key information in the log, generate accurate modeling query statements, and use visual components to visualize the data analysis results.

Benefits of technology

It quickly adapts to business data in different vertical fields, ensures the accuracy and efficiency of data processing, reduces the cost and threshold of model application, and improves the popularity of user experience and technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119537170B_ABST
    Figure CN119537170B_ABST
Patent Text Reader

Abstract

The present invention discloses a log modeling method based on semantic recommendation, belonging to the field of artificial intelligence technology. It includes extracting scenario description texts and log modeling statements from routine modeling work; classifying and organizing the extracted text content according to the modeling objectives and constructing question-and-answer pairs with special structures; integrating the text content required for modeling work and the constructed question-and-answer pairs into a knowledge base; the system uses large model and prompt word technologies to understand the modeling questions raised by users and infer the text content involved; using deep learning and semantic reasoning technologies, obtaining relevant question-and-answer pairs from the knowledge base and recommending similar question-and-answer pairs; the system automatically parses the recommended answers and their associated attributes, and modifies the content of the question-and-answer pairs according to specific constraint conditions in the user's question to generate adaptive answers. The present invention combines deep and semantic reasoning technologies to deeply analyze log data, automatically identify and understand key information in the logs, and generate accurate modeling query statements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, particularly to deep learning and natural language processing technologies, especially to large model technologies, and specifically to a log modeling method based on semantic recommendation. Background Art

[0002] Log modeling is a key technology in network security and data analysis. Its core lies in understanding and predicting the behavior of a system while detecting potential abnormal activities. This process begins with collecting the log data generated by the system, which can come from different sources such as servers, network devices, or applications. Then, features that can reflect the system state or behavior are extracted or constructed from the original log data. This may include time series data, IP addresses, operation types, response times, etc. The selection and construction of features directly affect the performance of the model. Statistical-based models, clustering models, or machine learning models are used to analyze and detect anomalies.

[0003] In the field of log modeling, existing technical solutions mainly focus on machine learning, natural language processing (NLP), and deep learning models, aiming to extract valuable information from massive log data for anomaly detection, trend analysis, and prediction.

[0004] Traditional machine learning methods, such as support vector machine (SVM), random forest (RF), etc., preprocess log data through feature engineering, convert it into computable feature vectors, and then use the model for classification or regression analysis to identify abnormal behaviors or predict system states. However, when dealing with unstructured or semi-structured log data, these methods often require a large amount of manual participation in feature selection, and the generalization ability of the model is limited.

[0005] The introduction of NLP technology, especially deep learning-based sequence models such as long short-term memory network (LSTM) and Transformer, can better understand and process the semantic information in logs. By learning the sequence features of log texts, it can automatically capture the correlation between logs and improve the accuracy of anomaly detection. At the same time, by calculating the semantic similarity of texts, it can assist in log clustering and anomaly detection, reducing manual intervention. However, when facing log data in vertical fields, due to the lack of in-depth understanding of business data, the accuracy and generalization ability of the model are limited, and a large amount of training data and computing resources are often required, resulting in high costs.

[0006] Based on the modeling and generation technology of general large models, although it can automate the generation of modeling statements to a certain extent and reduce the workload of data analysts, when this technology is applied in vertical fields, due to the lack of understanding of specific business scenarios, the generated modeling statements may not meet specific requirements, and the lack of intuitiveness and visual interaction increases the learning and usage threshold. To solve this problem, it is necessary to fine-tune the model or develop a dedicated domain model by combining domain knowledge, but this also increases the implementation complexity and cost.

[0007] In summary, models based on traditional machine learning and NLP, although capable of calculating text semantic similarity, often rely on manual intervention for confidence discrimination when dealing with complex log data in vertical fields, or the accuracy of the model is limited by the depth of knowledge in specific domains, resulting in insufficient generalization ability and adaptability. In addition, based on the modeling and generation technology of general large models, although it aims to automate the generation of modeling statements and reduce the burden of manual writing of models, due to the lack of in-depth understanding of specific business scenarios, the generated modeling statements may not accurately match actual needs, and their non-intuitiveness, high learning threshold, and lack of visual interaction affect the popularity of technology applications and user experience.

[0008] It can be seen that the development of log modeling technology is facing challenges such as how to improve the adaptability to specific business scenarios while maintaining the generalization ability of the model, and how to reduce the cost and threshold of model application. Future technology trends may focus more on the self-adaptability of the model, visual interaction, and the integration of domain knowledge to achieve more efficient and intelligent log analysis and processing. In view of this, the solution of the present invention proposes a log modeling method based on semantic recommendation. Summary of the Invention

[0009] The present invention aims to solve the problems in the existing log modeling methods, such as insufficient generalization ability and adaptability of the model, inability to accurately match actual needs, and affecting the popularity of technology applications and user experience. It proposes a log modeling method based on semantic recommendation, which deeply analyzes log data through deep learning technology combined with semantic reasoning algorithms, can quickly adapt to business data in different vertical fields, and ensure the accuracy and efficiency of data processing.

[0010] To achieve the above invention purpose, the technical solution of the present invention is as follows:

[0011] A log modeling method based on semantic recommendation, characterized by comprising the following steps:

[0012] Step a, extract text content including scenario description text and log modeling statements from regular modeling work;

[0013] Step b: Classify and organize the above text content according to the modeling objectives, and construct Q&A pairs with a special structure based on the classified text content;

[0014] Step c: Integrate the text content required for the modeling work and the constructed Q&A pairs into a knowledge base correspondingly;

[0015] Step d: When the user raises a specific modeling question, use the large model and prompt word technology to understand the question and infer the text content involved;

[0016] Step e: Use deep learning and semantic reasoning technology to obtain Q&A pairs related to the text content from the knowledge base, and recommend Q&A pairs similar to the question;

[0017] Step f: For the recommended Q&A pairs, the system automatically analyzes and reads their associated attributes, and renders them in combination with the visualization component; or use the large model and prompt word technology to modify the content of the Q&A pairs according to the specific constraints in the user's question, generate an adaptive answer and fill it into the visualization component to intuitively display the data processing results.

[0018] Further, the scenario description text and log modeling statements at least include descriptive text of the modeling target scenario, database description text, database field text, data field text, business requirement text, data analysis description text, and log modeling SQL statements.

[0019] Further, the data content included in the scenario description text and log modeling statements at least includes:

[0020] Modeling target description, including the business expected results, directions, and purposes to be achieved by the modeling;

[0021] Modeling SQL description, including the specific SQL statements written during the data analysis process and the statement remarks;

[0022] Modeling process description, including the understanding of data, scenarios, and the data calculation logic, judgment logic, and analysis logic used during the data modeling process;

[0023] Log data structure, including the description of the database and table structure for analyzing data during data modeling;

[0024] Modeling operator description, including operators or functions used to query and process data in the database;

[0025] Modeling reference text, including other documents and texts used to guide and assist the modeling during the modeling process.

[0026] Further, in step b, the question part of the question-and-answer pair is derived from the scenario description and the text of the modeling process description, and the answer part contains specific modeling statements or data processing logics for realizing the requirements.

[0027] Further, the answer part of the question-and-answer pair includes data visualization answer text and modeling visualization answer text. The data visualization answer text is used for data analysis to generate result data meeting the visualization requirements, and the modeling visualization answer text is used for data interaction in visual modeling.

[0028] Further, the question-and-answer pair with a special structure at least contains key / value pairs including questions, summaries, modeling processes, function operators, function operator methods, time, data types, data tables, fields used, modeling SQL, etc., which can enhance business understanding and semantic understanding.

[0029] Further, in step c, the text related to the database in the text content required for modeling is extracted, vectorized, and then constructed into a first knowledge base, and query and matching interfaces are provided; the text related to modeling in the text content required for modeling is extracted, vectorized, and then constructed into a second knowledge base, and query and matching interfaces are provided; the constructed question-and-answer pairs are extracted and constructed into a third knowledge base, and query and matching interfaces are provided; query interactions are provided among the first, second, and third knowledge bases.

[0030] Further, when the user raises a specific modeling question, the system first infers the log table fields in the database to which the question belongs, and then obtains the question-and-answer pairs related to the log table fields from the knowledge base; then, by screening out the associated text and question-and-answer pairs from the first, second, and third knowledge bases, they are matched with the user's question text, and the most similar and relevant question-and-answer pairs are output.

[0031] Further, in step f, the system automatically parses the answer and reads its associated attributes, and jointly renders with the visualization component, including: generating visualization SQL text, as well as fields and statistical data with time granularity according to the judgment of the user's question intention, while meeting the attribute constraints of the question text.

[0032] Further, in step f, using the large model and prompt word technology, the content of the question-and-answer pair is modified according to the specific constraint conditions in the user's question to generate an adaptive answer, including: generating modeling SQL text according to the judgment of the user's question intention, maintaining the modeling operators and modeling logics in the recommended question-and-answer, while meeting the attribute constraints of the question text.

[0033] In summary, the present invention has the following advantages:

[0034] 1. Through deep learning technology and combined with semantic reasoning algorithms, the present invention deeply analyzes log data, automatically identifies and understands key information in the logs, and generates accurate modeling query statements. At the same time, by using visualization components, complex data analysis results are presented in an intuitive chart form, reducing the learning threshold of data analysis and enhancing the user experience.

[0035] 2. The present invention utilizes semantic reasoning technology, adjusts it in combination with large models and prompt words, semantically understands user questions, automatically classifies modeling objectives, and automatically generates the required modeling statement text, achieving fully intelligent data analysis.

[0036] 3. Using the method of the present invention, there is no need to retrain the model, and it can quickly adapt to business data in different vertical fields, ensuring the accuracy and efficiency of data processing.

[0037] 4. The present invention constructs a knowledge base for business analysis, storing rich scenario modeling knowledge. By constructing a special data structure, the large model can understand the data and interface content in the business, realize the linkage with the visualization components, and provide a more intuitive and friendly user interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flowchart of a log modeling method recommended according to semantics provided by an embodiment of the present invention;

[0039] Figure 2 It is a schematic diagram of modeling data processing provided by an embodiment of the present invention;

[0040] Figure 3 It is a flowchart of question - answer pair construction provided by an embodiment of the present invention;

[0041] Figure 4 It is a flowchart of knowledge base construction provided by an embodiment of the present invention;

[0042] Figure 5 It is a flowchart of question - asking intention speculation and recommendation provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The following will describe the exemplary embodiments of the present disclosure in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and the examples provided are for clear description and easy understanding, and are not used to limit the protection scope of the present disclosure. For the sake of clarity and conciseness, the following description omits the description of well - known functions and explanatory descriptions.

[0044] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter. As used herein, the term "model" may represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on various technical solutions known currently and / or to be developed in the future.

[0045] As used herein, the term "model" can learn the corresponding association relationship between input and output from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes the input and provides the corresponding output by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms can be used interchangeably herein.

[0046] The explanations of some professional terms in this article are as follows:

[0047] Prompt: A prompt is to enhance the performance of a generative model in a specific knowledge field by inputting specific instructions or questions to the model. The structure of a prompt usually includes the following parts and can be flexibly combined according to task requirements: Context, Task Description, InputData, Output Format, Example.

[0048] SQL: The full name is Structured Query Language, that is, Structured Query Language, which is the standard language for managing relational databases. It is defined by the National Institute of Standards and Technology (NIST) of the United States and adopted by the International Organization for Standardization (ISO). It is the standard method for accessing and processing databases. The syntax and functions of SQL may vary slightly in different database management systems (such as Oracle, MySQL, SQL Server, PostgreSQL, etc.), but the core concepts and most of the syntax are common.

[0049] Data Definition Language (DDL): Used to create, modify, and delete database objects such as tables, views, indexes, etc. DDL commands include CREATE TABLE, CREATE INDEX, DROP TABLE, ALTER TABLE, etc. These commands allow database administrators to define the structure and layout of the database.

[0050] Data Manipulation Language (DML): Used to insert, update, and delete data. DML commands include INSERT, UPDATE, DELETE, which enable users to modify the data stored in the database.

[0051] Data Query Language (DQL): In fact, the data query language is usually regarded as part of DML, but it is mainly used to retrieve data. The most typical command is SELECT. The basic structure of the SELECT statement is indeed a query block composed of the SELECT clause, FROM clause, and WHERE clause, which is used to extract specific data sets from the database.

[0052] Data Control Language (DCL): Used to manage the security and permissions of the database, including granting or revoking user access rights to the database. DCL commands include GRANT and REVOKE, which are used to manage user roles and permissions to ensure data security and compliance. In addition, the ROLLBACK command also belongs to the transaction control language, which is used to roll back uncommitted transactions and return the database state to the state before the transaction started.

[0053] The present invention provides a log modeling method based on semantic recommendation. The core of the solution is to use semantic reasoning technology to semantically understand user questions, and then combine the constructed prompt words to guide the large model to generate the required modeling statement text to achieve a fully intelligent data analysis effect. This method does not require retraining the model, can quickly adapt to business data in different vertical fields, ensure the accuracy and efficiency of data processing. At the same time, by constructing a special data structure, the large model can understand the data and interface content in the business, realize the linkage with the visualization component, and provide a more intuitive and friendly user interaction experience.

[0054] Referring to the flowchart as Figure 1 shown, this method mainly includes the following steps:

[0055] Step 1: As Figure 1 shown in box 101 in , extract the scenario description text and log modeling statements from the regular modeling work. These information carry the business logic and data processing requirements, which are the data basis for constructing the special structure of the question and answer pairs later.

[0056] Among them, the scenario description text includes texts such as analysis scenarios, business requirements, and data analysis. The scenario description text can be obtained through various means, including submission, web crawling, downloading, database query, etc. The data forms are not limited to documents, texts, voices, videos, etc. Through existing technical means, various data forms can be converted into unified text content, which at least includes descriptive text of the modeling target scenario, database description text, database field text, data field text, business requirement text, data analysis description text, and log modeling SQL statements.

[0057] As Figure 2 shown in the schematic diagram for providing modeling data processing, box 100 in the figure represents obtaining scenario data for regular modeling work. In the regular data analysis process of various industries, when conducting modeling work on structured log data, the data required, produced, and referred to during the process are all scenario data.

[0058] Box 110 represents the data content that the obtained scenario data for regular modeling work at least includes: modeling target description 111, that is, the business expected results, directions, and purposes to be achieved by modeling (for example: analyzing file logs, counting the number of file downloads per day for drawing trend charts); modeling SQL description 112, that is, the specific SQL statements written during the data analysis process and the statement remarks; modeling process description 113, that is, during the data modeling process, the understanding of the data, the understanding of the scenario, and text contents such as the data calculation logic, judgment logic, and analysis logic used; log data structure 114, that is, the description of the database and table structure for analyzing data during data modeling, including data descriptions constructed by data query language (DQL), data manipulation language (DML), data definition language (DDL), and data control language (DCL); modeling operator description 115, that is, the operators or functions used to query and process data in the database. They can be used to extract, filter, transform, and summarize data to meet specific query requirements (for example: aggregation operators such as summation, maximum value, average value, etc.); modeling reference text 116, that is, other documents and texts used to guide and assist modeling during the modeling process.

[0059] Step 2: Refer to Figure 1 box 102 in to classify and organize the collected scenario description text and log modeling statements, and construct question-and-answer pairs with special structures, thereby improving the accuracy of semantic understanding and recommendation.

[0060] In this step, the question part of the question-answer pair usually comes from the scenario description and the text of the modeling process description, while the answer part contains specific modeling statements or data processing logics to meet the requirements. The question text in the question-answer pair asks questions around business problems (e.g., How many files were downloaded in the system today?). The answer text is a response structure text that responds to the questioned text, and it can be in structured text formats such as json or xml.

[0061] In this embodiment, the special structure of the question-answer pair adopts the key / value pair, which at least includes the question, summary, modeling process, function operator, function operator method, time, data type, data table, fields used, modeling SQL, and key / value pairs that can enhance business understanding and semantic understanding.

[0062] As Figure 3 shown is a flowchart for constructing the question-answer pair. In the figure, box 200 represents classifying the scenario of the modeling target. Since there are one or more business demands in the user's question, which can be generating statements that can be used to normally execute the modeling, recommending original modeling statements from the knowledge base, generating statements that are consistent with the attribute constraint conditions of the question, or generating modeling statements that can be used for user data visualization rendering, etc. Therefore, it is necessary to classify the user's demands by scenario to facilitate selecting the correct data structure when constructing the question-answer pair later.

[0063] Taking the example shown in box 210, according to the scenario classification result of the modeling target in box 200, combined with the modeling target description 111 and the modeling process description 113, construct the question in the question-answer pair. Since the user's question text may have incomplete expressions and unclear targets, the constructed question needs to cover common business problems of the user to facilitate subsequent recommendations that can be truly used in business production.

[0064] As shown in box 220, construct a structured answer according to the modeling SQL description 112. Specifically, based on the data content included in box 110, extract the SQL statement in the modeling SQL description 112. After being understood by the large model, automatically extract the data in the SQL, including the database name, table name, table alias, original fields, alias fields, field types, operator operations, etc., and at the same time generate the analysis logic text of the SQL.

[0065] As shown in boxes 230 and 240, according to the classification result of box 200, then construct the answer of box 220 into a data visualization answer text 230 and a modeling visualization answer text 240 respectively. The answer of box 230 can be used for data analysis to generate result data that meets the visualization requirements, and the answer of box 240 can be used for data interaction in visual modeling.

[0066] Step Three: Refer to Figure 1In the box 103, the text content used in business modeling work such as database field text and log modeling SQL statements, as well as the constructed question-and-answer pairs, are integrated into a knowledge base. The construction of the knowledge base adopts advanced information retrieval and natural language processing technologies, which can automatically identify and understand the structural characteristics of log tables in the database and the syntax and logic of modeling statements, and provide semantic matching and recommendation capabilities.

[0067] Specifically, taking the flowchart of the knowledge base construction as shown in Figure 4 as an example, the box 300 in the figure represents extracting the database-related text in the text content required for modeling (for example, extracting the database, table, and description text in the database field text), constructing the first knowledge base after vectorization processing, and providing query and matching interfaces; the box 310 represents extracting the modeling-related text in the text content required for modeling (for example, extracting the modeling reference text in the log modeling SQL statement), constructing the second knowledge base after vectorization processing, and providing query and matching interfaces; the box 320 represents extracting the question-and-answer pairs formed by the boxes 210, 230, and 240, constructing the third knowledge base, and providing query and matching interfaces; the box 330 represents managing the first, second, and third knowledge bases and providing query interaction. In an alternative embodiment, an NLP natural language processing model is used to perform vectorization processing on the text.

[0068] Step Four: Referring to Figure 1 the box 104, when the user asks a specific modeling question, based on the modeling question text, first speculate on the log table fields in the database to which the question belongs.

[0069] Specifically, according to the user's question, using large model and prompt word technologies, understand and speculate on the question text (for example: How many files were downloaded by the system today?), and speculate on the log table fields in the database involved in the question (for example: file log table).

[0070] Step Five: Referring to Figure 1 the box 105, obtain the question-and-answer pairs corresponding to the log table fields to which the question belongs from the knowledge base, and recommend question-and-answer pairs similar or close to the question.

[0071] In this step, deep learning and semantic reasoning technologies are used to intelligently recommend question-and-answer pairs similar or close to the question. Specifically, when implementing, analyze the log table fields in the database associated with the data according to the question text, and screen out the associated text and question-and-answer pairs from the first knowledge base, the second knowledge base, and the third knowledge base, and match them with the question text, so as to output the most similar and close question-and-answer pairs.

[0072] Step Six: Referring to Figure 1 the box 106, for the recommended question-and-answer pairs, the system can automatically parse the associated attributes in the answer and perform rendering in combination with the visualization component; or asFigure 1 As shown in box 107 in

[0073] In this step, since there is one or more business requirements in the user's question, it is necessary to analyze the intention of the user's question, judge its business type and generate a correct answer in order to output the modeling SQL structured data that meets the question intention.

[0074] Such as Figure 5 As shown is the flowchart for inferring and recommending question intentions provided by this embodiment. As shown in box 500, first analyze the intention of the question and judge the business type.

[0075] In an alternative embodiment, according to the judgment, visual SQL text as shown in box 501, fields with time granularity, and statistical data can be generated, while meeting the attribute constraints of the question text. It can be understood that the generated SQL text can automatically generate new fields with time granularity to meet the requirements for each time scale in visualization, and generate new statistical fields using new operators to meet the statistical data under the time granularity.

[0076] In an alternative embodiment, according to the judgment, modeling SQL text as shown in box 502 can be generated, maintaining the modeling operators and modeling logic in the recommended Q&A, while meeting the attribute constraints of the question text. It can be understood that the data in the recommended Q&A comes from the constructed knowledge base, and the data is all constructed from the obtained business data. Therefore, it can be determined that the modeling operators and modeling logic included in the data are correct and can be used for business analysis. Therefore, when generating the modeling SQL text, it is only necessary to maintain the modeling operators and logic in the original Q&A and use the large model to only modify the attribute constraints (for example: modify the time attribute constraint, modify the field value size constraint).

[0077] Step seven, referring to Figure 1 As shown in box 108 in

[0078] Specifically, as Figure 5 As shown in box 503 in

[0079] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Any simple modification or equivalent change made to the above embodiments based on the technical essence of the present invention shall fall within the protection scope of the present invention.

Claims

1. A log modeling method based on semantic recommendation, characterized in that: The steps include: Step a, extracting text content including scenario description text and log modeling statements from conventional modeling work; Step b: classify and organize the above text content according to the modeling goal, and construct a special structured question-answer pair based on the classified text content; the special structured question-answer pair at least includes question, summary, modeling process, function operator, function operator method, time, data type, data table, used fields, modeling SQL and key / value pairs that can enhance business understanding and semantic understanding; Step c: Integrate the text content required for modeling and the constructed question-answer pairs into a knowledge base; Step d: When the user raises a specific modeling question, the large model and prompt word technology are used to understand the question and infer the text content involved; Step e: Using deep learning and semantic reasoning technology, obtain question-answer pairs related to the text content from the knowledge base, and recommend question-answer pairs that are similar or close to the question; Step f: For the recommended question-answer pairs, the system automatically parses the answers and reads their associated attributes, and renders them in conjunction with the visualization component; or uses the big model and prompt word technology to modify the question-answer pair content according to the specific constraints in the user's question, generate adaptive answers and fill them into the visualization component to intuitively display the data processing results; The scenario description text and log modeling statement at least include modeling target scenario description text, database description text, database field text, data field text, business requirement text, data analysis description text and log modeling SQL statement; the data content contained in the scenario description text and log modeling statement at least includes: Description of modeling objectives, including expected business results, directions, and purposes that modeling needs to achieve; Modeling SQL description, including specific SQL statements written during data analysis and statement comments; Description of the modeling process, including the understanding of data and scenarios during the data modeling process, as well as the data calculation logic, judgment logic, and analysis logic used; Log data structure, including the description of the database and table structure of analytical data during data modeling; Modeling operator descriptions, including operators or functions used to query and process data in the database; Modeling reference texts include other documents and texts used to guide and assist modeling during the modeling process.

2. The log modeling method based on semantic recommendation according to claim 1, characterized in that: In step b, the question part of the question-answer pair comes from the scenario description and modeling process description text, and the answer part contains the specific modeling statements or data processing logic to implement the requirements.

3. The log modeling method based on semantic recommendation as claimed in claim 2, characterized in that: The answer part of the question and answer pair includes data visualization answer text and modeling visualization answer text. The data visualization answer text is used for data analysis to generate result data that meets visualization requirements, and the modeling visualization answer text is used for data interaction of visualization modeling.

4. The log modeling method based on semantic recommendation according to claim 1, characterized in that: In step c, the text related to the database is extracted from the text content required for modeling, and is constructed into a first knowledge base after vectorization processing, and a query and matching interface is provided; the text related to modeling is extracted from the text content required for modeling, and is constructed into a second knowledge base after vectorization processing, and a query and matching interface is provided; the constructed question and answer pairs are extracted, constructed into a third knowledge base, and a query and matching interface is provided; query interaction is provided between the first, second, and third knowledge bases.

5. The log modeling method based on semantic recommendation as claimed in claim 4, characterized in that: When a user asks a specific modeling question, the system first infers the log table field in the database to which the question belongs, and then obtains the question-answer pairs related to the log table field from the knowledge base; then, it filters out related texts and question-answer pairs from the first, second, and third knowledge bases, matches them with the user's question text, and outputs the most similar and close question-answer pairs.

6. The log modeling method based on semantic recommendation according to claim 1, characterized in that: In step f, the system automatically parses the answer and reads its associated attributes, and renders it in conjunction with the visualization component, including: generating visual SQL text based on the judgment of the user's question intention, as well as fields and statistical data with time granularity, while satisfying the attribute constraints of the question text.

7. The log modeling method based on semantic recommendation according to claim 1, characterized in that: In step f, the large model and prompt word technology are used to modify the content of the question and answer pair according to the specific constraints in the user's question to generate an adaptive answer, including: generating a modeling SQL text based on the judgment of the user's question intention, maintaining the modeling operators and modeling logic in the recommended question and answer, and satisfying the attribute constraints of the question text.

Citation Information

Patent Citations

  • Knowledge question-answering method and device, equipment and storage medium

    CN116737908A

  • Question and answer pair generation method, computer equipment and storage medium

    CN117633190A