Text analysis method and device based on large model and electronic equipment

Through the big model splitting text analysis task, obtaining the target data set and generating BI reports, solving the problems of low retrieval efficiency and low accuracy in the existing technology, and realizing automated and efficient text analysis.

CN120256595APending Publication Date: 2025-07-04BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510314424.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, when users perform text analysis in a database, the search efficiency is low and the accuracy is low, and manual participation is required to lead to many cumbersome steps.

Method used

By obtaining the information type of input information, using a large model to split the text analysis task, obtaining the target data set and generating BI reports, reducing manual participation and improving retrieval efficiency and accuracy.

Benefits of technology

In the case of high complexity input information, text analysis tasks can be automatically split, required data tables can be obtained, accurate BI reports can be generated, manual search steps can be reduced, and text analysis efficiency and accuracy can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256595A_ABST
    Figure CN120256595A_ABST
Patent Text Reader

Abstract

The invention provides a text analysis method and device based on a large model and electronic device.The text analysis method based on the large model comprises the steps that an information type corresponding to input information is obtained; when the information type is a first information type, querying in a BI analysis data table library according to the input information to obtain a first target data set; executing a natural language query operation by adopting task planning information in a task planning information set corresponding to the first target data set to obtain a first query data set; and inputting the task planning information set and the first query data set into a large model for analysis, obtaining an analysis result, and generating a BI report corresponding to the input information according to the analysis result, thereby solving the problems of low retrieval efficiency and low accuracy caused by manual retrieval for the input information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a text analysis method, apparatus, and electronic device based on a large model. Background Art

[0002] With the development of science and technology, electronic devices can provide users with more and more services, improving the convenience of users' lives. For example, through the questions input by users, data can be directly retrieved from the database, and the content that the user wants to retrieve can be output. Among them, question-based data analysis refers to a technology for data analysis through dialogue. It uses technologies such as natural language understanding and knowledge graphs to help users quickly and conveniently search for data, interpret data, and visualize data. Among them, for example, simple question-based data analysis operations can be performed manually in the database, making the retrieval steps cumbersome and the retrieval efficiency and accuracy low. Summary of the Invention

[0003] This application aims to solve at least one of the technical problems in the related art to some extent.

[0004] To this end, the first object of this application is to propose a text analysis method based on a large model, which can split text analysis tasks, improve the accuracy of obtaining BI reports, and can reduce the cumbersome steps of manual retrieval without manual participation, thereby improving the efficiency and accuracy of text analysis.

[0005] The second object of this application is to propose a text analysis apparatus based on a large model.

[0006] The third object of this application is to propose an electronic device.

[0007] The fourth object of this application is to propose a computer-readable storage medium.

[0008] The fifth object of this application is to propose a computer program product.

[0009] To achieve the above object, the first aspect embodiment of this application proposes a text analysis method based on a large model, including the following steps:

[0010] Obtain the information type corresponding to the input information;

[0011] When the information type is the first information type, query in the BI analysis data table library according to the input information to obtain a first target data set, where the first information type is used to indicate that the complexity corresponding to the input information is greater than the first complexity threshold, and the similarity between each target data in the first target data set and the input information is less than the first similarity threshold and greater than the second similarity threshold, and the first similarity threshold is greater than the second similarity threshold;

[0012] Perform a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set to obtain a first query data set;

[0013] Input the task planning information set and the first query data set into a large model for analysis, obtain an analysis result, and generate a BI report corresponding to the input information according to the analysis result.

[0014] To achieve the above object, an embodiment of the second aspect of the present application proposes a text analysis device based on a large model, including:

[0015] A type acquisition unit for acquiring the information type corresponding to the input information;

[0016] A set acquisition unit for querying in a BI analysis data table library according to the input information when the information type is the first information type to obtain a first target data set, where the first information type is used to indicate that the complexity corresponding to the input information is greater than a first complexity threshold, and the similarity between each target data in the first target data set and the input information is less than a first similarity threshold and greater than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold;

[0017] The set acquisition unit is further configured to perform a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set to obtain a first query data set;

[0018] A report generation unit for inputting the task planning information set and the first query data set into a large model for analysis, obtaining an analysis result, and generating a BI report corresponding to the input information according to the analysis result.

[0019] To achieve the above object, an embodiment of the third aspect of the present application proposes an electronic device, including: a processor and a memory communicatively connected to the processor;

[0020] The memory stores computer execution instructions;

[0021] The processor executes the computer execution instructions stored in the memory to implement the method according to any one of the above first aspects.

[0022] To achieve the above object, an embodiment of the fourth aspect of the present application proposes a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of the above first aspects.

[0023] To achieve the above object, an embodiment of the fifth aspect of the present application provides a computer program product, including a computer program, which when executed by a processor implements the method according to any one of the above first aspect.

[0024] The text analysis method, device and electronic device based on a large model provided by the present application obtain the information type corresponding to the input information; when the information type is the first information type, query in the BI analysis data table library according to the input information to obtain a first target data set, where the first information type is used to indicate that the complexity corresponding to the input information is greater than the first complexity threshold, the similarity between each target data in the first target data set and the input information is less than the first similarity threshold and greater than the second similarity threshold, and the first similarity threshold is greater than the second similarity threshold; perform a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set to obtain a first query data set; input the task planning information set and the first query data set into the large model for analysis to obtain an analysis result, and generate a BI report corresponding to the input information according to the analysis result, which solves the problem that manual retrieval is required for the input information, resulting in low retrieval efficiency and low accuracy. Since the information type of the input information can be judged, the task planning information of the input information can be obtained when the complexity of the input information is relatively high, the text analysis task can be split, and the target data set corresponding to the input information can be obtained, the required data table corresponding to the input information can be obtained, and it is not necessary to analyze all of them, which can improve the accuracy of obtaining the BI report, and there is no need for manual participation in retrieval, which can reduce the cumbersome steps of manual retrieval and improve the efficiency and accuracy of text analysis.

[0025] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:

[0027] Figure 1 is a schematic flowchart of a text analysis method based on a large model provided by an embodiment of the present application;

[0028] Figure 2 is a schematic flowchart of a text analysis method based on a large model provided by an embodiment of the present application;

[0029] Figure 3Schematic diagram for an example of a text analysis method based on a large model provided by an embodiment of the present application;

[0030] and Figure 4 Schematic diagram of the structure of a text analysis device based on a large model provided by an embodiment of the present application. Detailed implementation manners

[0031] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.

[0032] The text analysis method and device based on a large model according to an embodiment of the present application will be described below with reference to the accompanying drawings.

[0033] Figure 1 Schematic flowchart of a text analysis method based on a large model provided by an embodiment of the present application.

[0034] To address this problem, the embodiments of the present application provide a text analysis method based on a large model, which can split text analysis tasks, improve the accuracy of obtaining BI reports, and eliminate the need for manual retrieval, reducing the cumbersome steps of manual retrieval and improving the efficiency and accuracy of text analysis. As Figure 1 shown, the text analysis method based on a large model includes the following steps:

[0035] Step 101, obtain the information type corresponding to the input information;

[0036] According to some embodiments, the execution subject of the embodiments of the present application may be, for example, an electronic device. This electronic device does not specifically refer to a certain fixed device, and the name of this electronic device is not limited. This electronic device may also be referred to as a terminal, a mobile device, etc. For example, when the device identifier of this electronic device changes, this electronic device may also change accordingly.

[0037] In some embodiments, the input information may be, for example, the information received during the question asking, and this input information may be the question received during the question asking operation. This input information does not specifically refer to a certain fixed information. For example, when the acquisition method of the input information changes, this input information may also change accordingly. For example, when the acquisition time point corresponding to the input information changes, this input information may also change accordingly. Among them, the acquisition method of the input information is not limited. For example, this input information may be input by clicking through the input method on the electronic device, or may be voice information input through the voice input control.

[0038] According to some embodiments, the information type can be used, for example, to refer to the type to which the input information belongs. Among them, an input information can correspond to one information type, different input information can correspond to different information types, and different input information can also correspond to the same information type. The embodiments of the present application do not limit this. Among them, the preset information type is not limited. For example, when the number of types corresponding to the preset information type changes, the preset information type can also change accordingly. For example, when the specific type corresponding to the preset information type changes, the preset information type can also change accordingly.

[0039] In some embodiments, there is no limitation on the method for obtaining the information type. For example, it can be related to the current application scenario, for example, it can also be related to the instruction input by the user, and for example, it can also be related to the amount of information of the input information.

[0040] In some embodiments, the information type corresponding to the input information can be obtained.

[0041] Step 102, when the information type is the first information type, query in the BI analysis data table library according to the input information to obtain a first target data set, where the first information type is used to indicate that the complexity corresponding to the input information is greater than a first complexity threshold, and the similarity between each target data in the first target data set and the input information is less than a first similarity threshold and greater than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold;

[0042] In some embodiments, the first in the first information type can be used, for example, to distinguish it from other information types. The first information type is used to indicate that the complexity corresponding to the input information is greater than a first complexity threshold, and can be used, for example, to indicate that the complexity of the input information is relatively complex.

[0043] According to some embodiments, the complexity can be used, for example, to indicate the complexity level of the input information. This complexity does not specifically refer to a certain fixed value. For example, when the input information changes, the complexity of the input information can also change accordingly. For example, when the way to determine the complexity of the input information changes, the complexity of the input information can also change accordingly.

[0044] According to some embodiments, the complexity threshold can be a threshold related to the complexity. The complexity threshold can correspond to the current application scenario information, for example, and can also be determined according to the threshold setting instruction. The embodiments of the present application do not limit the determination method of the complexity threshold.

[0045] In some embodiments, the first in the first complexity threshold, for example, can be distinguished from the remaining complexity thresholds. This first complexity threshold does not specifically refer to a certain fixed threshold. For example, when a modification instruction for the first complexity threshold is received, the first complexity threshold can also change accordingly. For example, when determining the first complexity threshold based on application scenario information, when the application scenario information changes, the first complexity threshold can also change accordingly.

[0046] In some embodiments, the BI analysis data table library, for example, can be a collective formed by converging at least one data table. This BI analysis data table library does not specifically refer to a certain fixed data table library. For example, when the data tables included in the BI analysis data table library change, the BI analysis database can also change accordingly. For example, when the number of data tables in the BI analysis data table library changes, the BI analysis database can also change accordingly.

[0047] According to some embodiments, the target data set, for example, can be a set formed by converging at least one target data. This target data set does not specifically refer to a certain fixed set. For example, when the query method changes, the target data set can also change accordingly. For example, when the number of data included in the target data set changes, the target data set can also change accordingly.

[0048] In some embodiments, the similarity between each target data in the first target data set and the input information is less than the first similarity threshold and greater than the second similarity threshold, and the first similarity threshold is greater than the second similarity threshold. The similarity can be used to indicate the degree of similarity between each data in the data table library and the input information. Among them, the first in the first similarity threshold is used to distinguish it from the remaining similarities.

[0049] In some embodiments, there is no limitation on the value-taking method for the first similarity threshold and the second similarity threshold. For example, it can be determined according to application scenario information, for example, it can also be set according to a setting instruction, or it can be set by a combination of multiple methods. The embodiments of the present application do not limit this.

[0050] In some embodiments, when the information type is the first information type, that is, when the complexity of the input information is greater than the first complexity threshold, querying in the BI analysis data table library according to the input information can obtain data whose similarity with the input information is less than the first similarity threshold and greater than the second similarity threshold, and add this data to the first target data set. That is, the first target data set can be obtained, and the similarity between each target data in the first target data set and the input information is less than the first similarity threshold and greater than the second similarity threshold, and the first similarity threshold is greater than the second similarity threshold.

[0051] Step 103: Perform a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set to obtain a first query data set;

[0052] In some embodiments, the task planning information can be used, for example, to indicate the processing method of the input information. For example, it can include how to ask questions about the input information. Among them, the acquisition method of the task planning information is not limited. For example, it can be generated by a large model. When the large model changes, the task planning information can also change accordingly.

[0053] According to some embodiments, the task planning information set can be, for example, a collective formed by at least one task planning information. The task planning information set does not specifically refer to a certain fixed set. For example, when a certain task planning information in the task planning information set changes, the task planning information set can also change accordingly. For example, when the number of task planning information corresponding to the task planning information set changes, the task planning information set can also change accordingly.

[0054] In some embodiments, the first in the first query data set is used to distinguish it from the rest of the query data sets. The first query data set can be, for example, a collective formed by at least one query data. The first query data set does not specifically refer to a certain fixed set. For example, when the amount of query data included in the first query data set changes, the first query data set can also change accordingly. For example, when the number of operation times during the natural language query operation changes, the first query data set can also change accordingly.

[0055] In some embodiments, a natural language query operation can be performed using each task planning information in the task planning information set corresponding to the first target data set to obtain a first query data set.

[0056] Step 104: Input the task planning information set and the first query data set into a large model for analysis to obtain an analysis result, and generate a BI report corresponding to the input information according to the analysis result.

[0057] In some embodiments, the large model can be, for example, a model that has been trained and can be used in the text analysis process. Among them, the large model can perform multiple operations. Specifically, for example, different sub-models in the large model can perform operations. It can also be multiple large models, and different large models perform different operations. The embodiments of the present application do not limit this.

[0058] In some embodiments, the analysis result can be, for example, the result obtained by data analysis.

[0059] According to some embodiments, BI can be, for example, a technology and method that utilizes modern information technology to collect, manage, and analyze structured and unstructured business data and information, create and accumulate business knowledge and insights, so as to improve the level of business decision-making, take effective business actions, improve business processes, and enhance business performance.

[0060] Among some embodiments, for example, the task planning information set and the first query data set can be input into a large model for analysis to obtain an analysis result, and based on the analysis result, a BI report corresponding to the input information is generated. Herein, the name of the BI report is not limited. For example, it can also be called a BI analysis report, a BI generation report, etc.

[0061] The text analysis method, device, and electronic device based on a large model provided in this application obtain the information type corresponding to the input information; when the information type is the first information type, query in the BI analysis data table library according to the input information to obtain a first target data set, where the first information type is used to indicate that the complexity corresponding to the input information is greater than a first complexity threshold, the similarity between each target data in the first target data set and the input information is less than a first similarity threshold and greater than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold; perform a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set to obtain a first query data set; input the task planning information set and the first query data set into the large model for analysis to obtain an analysis result, and based on the analysis result, generate a BI report corresponding to the input information, which solves the problem that manual retrieval is required for the input information, resulting in low retrieval efficiency and low accuracy. Since the information type of the input information can be judged, the task planning information of the input information can be obtained when the complexity of the input information is high, the text analysis task can be split, and the target data set corresponding to the input information can be obtained, the required data table corresponding to the input information can be obtained, and it is not necessary to analyze all of them, which can improve the accuracy of obtaining the BI report. Moreover, since there is no need for manual participation in retrieval, the cumbersome steps of manual retrieval can be reduced, and the efficiency and accuracy of text analysis can be improved. In addition, the situation of asking questions multiple times for multiple BI systems can be reduced, the text analysis steps can be reduced, and the text analysis efficiency can be improved.

[0062] This embodiment provides another text analysis method based on a large model. Figure 2 It is a schematic flowchart of a text analysis method based on a large model provided by an embodiment of this application.

[0063] As Figure 2As shown, the large model-based text analysis method may include the following steps:

[0064] Step 201, obtaining the information type corresponding to the input information;

[0065] The specific process is as described above and will not be elaborated here.

[0066] According to some embodiments, Figure 3 A schematic example diagram showing a large model-based text analysis method is shown. As Figure 3 shown, when the input text is obtained, a problem routing can be performed first, that is, the information type corresponding to the input information can be determined. Specifically, it may include: determining the information type of the input information according to the determination priority of the information type.

[0067] According to some embodiments, the obtaining of the problem type corresponding to the input text includes:

[0068] When a type selection instruction of the display interface is detected, according to the type selection instruction, determining the problem type corresponding to the input text;

[0069] Or,

[0070] When the type selection instruction is not received, obtaining a first vector corresponding to the input text;

[0071] Using the first vector to perform multiple rounds of retrieval in the problem vector library through table routing to obtain a retrieval result set, and obtaining the problem type corresponding to the input text according to the retrieval result set, where the retrieval result set includes at least one retrieval result, and the similarity between each retrieval result in the at least one retrieval result and the first vector is greater than the similarity threshold.

[0072] According to some embodiments, for example, an analysis model control can be displayed on the display interface of an electronic device, and a selection instruction of the information type can be obtained through the high analysis model control, and the electronic device can determine the information type according to the selection instruction.

[0073] In some embodiments, when no selection instruction is obtained, the vector of the input information can be retrieved to obtain the information type corresponding to the input information. Specifically, for example, it can be retrieved in the problem vector library to obtain at least one retrieval result whose similarity degree with the input information is greater than the similarity threshold. Among them, the retrieval result can be, for example, historical input information. The information type corresponding to the input information can be determined according to the information type corresponding to the historical input information.

[0074] According to some embodiments, the obtaining of the problem type corresponding to the input text includes:

[0075] In the case where the type selection instruction is not received and the retrieval result set is not obtained, the input text is input into the large model for classification processing to obtain the problem type corresponding to the input text. Therefore, in the case where information types are not obtained through retrieval, classification can be performed according to the large model to obtain the information type corresponding to the input information. Among them, the large model can be, for example, a model that has been trained and can be used for classification processing. That is, the information type of the current input information can be obtained through the information type of historical input information, which can improve the accuracy of information type determination.

[0076] Step 202, in the case where the information type is the first information type, query in the BI analysis data table library according to the input information to obtain a first target data set, where the first information type is used to indicate that the complexity corresponding to the input information is greater than a first complexity threshold, and the similarity between each target data in the first target data set and the input information is less than a first similarity threshold and greater than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold;

[0077] The specific process is as described above and will not be elaborated here.

[0078] According to some embodiments, in the case where the information type of the input information is the first information type, complex question numbers can be executed.

[0079] Among them, the BI analysis data table library can include multiple data tables, and the number of data tables can be greater than a data table quantity threshold. That is, there can be hundreds of data tables in the BI analysis data table library.

[0080] According to some embodiments, the method further includes:

[0081] In the case where the problem type is the first information type, query in the BI analysis data table library according to the input text to obtain a second target data set, and use the second target data set as the analysis result corresponding to the input information, where the similarity between each target data in the second target data set and the input text is greater than the first similarity threshold.

[0082] In some embodiments, when the problem type is the first information type, that is, the information type of the input information is a simple question number type, an analysis report corresponding to the input information can be obtained through a corresponding analysis method. Among them, simple question numbers can be used to indicate natural language query operations on input information.

[0083] In some embodiments, for example, the first similarity threshold can be 0.97, and the second similarity threshold can be 0.8. Therefore, the second target data set can include at least one historical analysis result with a relatively high similarity. This data set can be used as the analysis result without further analysis, saving resources.

[0084] In some embodiments, for example, the first similarity threshold can be 0.97, and the second similarity threshold can be 0.8. Therefore, the first target data set can include at least one historical analysis result with a similarity between 0.8 and 0.97. The first target data set can be input into a large model for reference to generate a set of task planning information corresponding to the input information. The set of task planning information can be, for example, the information obtained after splitting the input information into tasks. This task planning information can be used to indicate the tasks required for the input information. Among them, the acquisition method of the set of task planning information is not limited. It can be determined according to the information length of the input information or according to the reference scenario information corresponding to the input information.

[0085] Step 203, perform a natural language query operation using each task planning information in the set of task planning information corresponding to the first target data set to obtain a first query data set;

[0086] The specific process is as described above and will not be elaborated here.

[0087] In some embodiments, for example, the set of task planning information corresponding to the current first target data set can be obtained through historical planning task information. Or the set of task planning information can be obtained through a large model.

[0088] According to some embodiments, the performing a natural language query operation using each task planning information in the set of task planning information corresponding to the first target data set to obtain a first query data set includes:

[0089] Obtain the application scenario information corresponding to the input text;

[0090] Obtain at least one structured query language (SQL) data table library related to the application scenario;

[0091] According to each task planning information in the set of task planning information corresponding to the first target data set, perform a natural language query operation in each SQL data table library in the at least one SQL data table library to obtain a first query data set.

[0092] In some embodiments, a simple question count can be performed on each task planning information to obtain the query data corresponding to the SQL data table library.

[0093] Step 204: Input the task planning information set and the first query data set into the large model for analysis, obtain the analysis result, and generate a BI report corresponding to the input information according to the analysis result.

[0094] The specific process is as described above and will not be elaborated here.

[0095] According to some embodiments, the inputting the task planning information set and the first query data set into the large model for analysis and obtaining the analysis result includes:

[0096] Input the task planning information set and the query data set into the large model for code generation operation to obtain data analysis code.

[0097] Execute the data for data analysis to obtain the analysis result.

[0098] According to some embodiments, the BI report may also be referred to as a BI analysis summary report, for example.

[0099] Step 205: In the case where the information type is the second information type, retrieve in the milvus data table library according to the input text to obtain a data table identifier set corresponding to the input text, and obtain the table structure of the data tables corresponding to each data table identifier in the data table identifier set, where the second information type is used to indicate that the complexity corresponding to the input text is less than a second complexity threshold, and the second complexity threshold is less than the first complexity threshold.

[0100] The related process is as described above and will not be elaborated here.

[0101] In some embodiments, the data table identifier set may include at least one data table identifier. The table structure of the data table may be referred to as DDL, for example. The table structure of the data table includes the field name, field type, data table content, etc. of the data table.

[0102] Step 206: Use an intent recognition model to recognize the input text and obtain the three elements corresponding to the input text.

[0103] In some embodiments, the three elements may include an entity, an indicator, and a time, for example.

[0104] Step 207: Input the table structure of the data tables corresponding to each data table identifier and the three elements corresponding to the input text into the large model to generate a Structured Query Language (SQL).

[0105] Step 208, query in the data table library according to the Structured Query Language (SQL), obtain a second query data set, and use the second query data set as the analysis result corresponding to the input text.

[0106] In one embodiment, when the information type is the second information type, retrieve in the milvus data table library according to the input text, obtain a data table identifier set corresponding to the input text, and obtain the table structures of the data tables corresponding to each data table identifier in the data table identifier set. Use an intent recognition model to identify the input text and obtain three elements corresponding to the input text; input the table structures of the data tables corresponding to each data table identifier and the three elements corresponding to the input text into the large model to generate a Structured Query Language (SQL); query in the data table library according to the Structured Query Language (SQL), obtain a second query data set, and use the second query data set as the analysis result corresponding to the input text. Therefore, when performing a simple retrieval, retrieval can be performed through the three elements and the table structure of the data table, which can improve the retrieval efficiency and accuracy.

[0107] To implement the above embodiments, the present application also proposes a text analysis device based on a large model.

[0108] Figure 4 It is a schematic structural diagram of a text analysis device based on a large model provided by an embodiment of the present application.

[0109] As Figure 4 shown, the text analysis device based on a large model includes:

[0110] A type acquisition unit 401, configured to acquire an information type corresponding to input information;

[0111] A set acquisition unit 402, configured to query in the BI analysis data table library according to the input information when the information type is the first information type, and obtain a first target data set, where the first information type is used to indicate that the complexity corresponding to the input information is greater than a first complexity threshold, and the similarity between each target data in the first target data set and the input information is less than a first similarity threshold and greater than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold;

[0112] The set acquisition unit 402 is further configured to perform a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set, and obtain a first query data set;

[0113] A report generation unit 403 is configured to input the task planning information set and the first query data set into a large model for analysis, obtain an analysis result, and generate a BI report corresponding to the input information according to the analysis result.

[0114] According to some embodiments, the report generation unit 403 is further configured to:

[0115] In the case where the information type is the second information type, retrieve in the milvus data table library according to the input information, obtain a data table identifier set corresponding to the input information, and obtain the table structures of the data tables corresponding to the data table identifiers in the data table identifier set, where the second information type is used to indicate that the complexity corresponding to the input information is less than a second complexity threshold, and the second complexity threshold is less than the first complexity threshold;

[0116] Adopt an intention recognition model to recognize the input information and obtain three elements corresponding to the input information;

[0117] Input the table structures of the data tables corresponding to the data table identifiers and the three elements corresponding to the input information into the large model to generate a Structured Query Language (SQL);

[0118] Query in the data table library according to the Structured Query Language (SQL), obtain a second query data set, and use the second query data set as the analysis result corresponding to the input information.

[0119] According to some embodiments, when the type acquisition unit 401 is configured to acquire the information type corresponding to the input information, it is specifically configured to:

[0120] In the case where a type selection instruction of the display interface is detected, determine the information type corresponding to the input information according to the type selection instruction;

[0121] Or,

[0122] In the case where the type selection instruction is not received, obtain a first vector corresponding to the input information;

[0123] Perform multiple rounds of retrieval in the problem vector library using the first vector through table routing to obtain a retrieval result set, and obtain the information type corresponding to the input information according to the retrieval result set, where the retrieval result set includes at least one retrieval result, and the similarity between each retrieval result in the at least one retrieval result and the first vector is greater than a similarity threshold.

[0124] According to some embodiments, when the type acquisition unit 401 is configured to acquire the information type corresponding to the input information, it is specifically configured to:

[0125] In the case where the type selection instruction is not received and the retrieval result set is not obtained, the input information is input into the large model for classification processing to obtain the information type corresponding to the input information.

[0126] According to some embodiments, the report generation unit 403 is further configured to:

[0127] In the case where the information type is the first information type, query in the BI analysis data table library according to the input information to obtain a second target data set, and use the second target data set as the analysis result corresponding to the input information, where the similarity between each target data in the second target data set and the input information is greater than the first similarity threshold.

[0128] According to some embodiments, when the report generation unit 403 is configured to input the task planning information set and the first query data set into the large model for analysis to obtain an analysis result, it is specifically configured to:

[0129] Input the task planning information set and the query data set into the large model for code generation operation to obtain data analysis code;

[0130] Execute the data for data analysis to obtain an analysis result.

[0131] According to some embodiments, when the set acquisition unit 402 is configured to perform a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set to obtain a first query data set, it is specifically configured to:

[0132] Obtain the application scenario information corresponding to the input information;

[0133] Obtain at least one structured query language (SQL) data table library related to the application scenario;

[0134] According to each task planning information in the task planning information set corresponding to the first target data set, perform a natural language query operation in each structured query language (SQL) data table in the at least one structured query language (SQL) data table library to obtain a first query data set.

[0135] It should be noted that the foregoing explanation of the embodiments of the text analysis method based on the large model also applies to the text analysis device based on the large model in this embodiment, and will not be repeated here.

[0136] The text analysis device based on a large model provided by this application includes a type acquisition unit for acquiring the information type corresponding to the input information; a set acquisition unit for querying in the BI analysis data table library according to the input information when the information type is the first information type, and acquiring a first target data set, where the first information type is used to indicate that the complexity corresponding to the input information is greater than a first complexity threshold, and the similarity between each target data in the first target data set and the input information is less than a first similarity threshold and greater than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold; the set acquisition unit is further configured to perform a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set to acquire a first query data set; a report generation unit for inputting the task planning information set and the first query data set into the large model for analysis to obtain an analysis result, and generating a BI report corresponding to the input information according to the analysis result, which solves the problem that manual retrieval is required for the input information, resulting in low retrieval efficiency and low accuracy. Since the information type of the input information can be judged, the task planning information of the input information can be obtained when the complexity of the input information is relatively high, the text analysis task can be split, the accuracy of obtaining the BI report can be improved, and manual participation in retrieval is not required, which can reduce the cumbersome steps of manual retrieval and improve the efficiency and accuracy of text analysis.

[0137] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0138] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer execution instructions, which are used to implement the method provided in the foregoing embodiments when executed by a processor.

[0139] To implement the above embodiments, this application also proposes a computer program product including a computer program, which implements the method provided in the foregoing embodiments when executed by a processor.

[0140] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0141] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legitimate uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing the relevant user information before the user uses the function. In addition, any necessary steps should be taken to defend and safeguard access to such personal information data and ensure that others with access to the personal information data comply with its privacy policies and procedures.

[0142] This application is expected to provide an implementation for users to selectively block the use or access to personal information data. That is, this application is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of the user.

[0143] In the description of the foregoing embodiments, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0144] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0145] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred implementation of this application includes additional implementations, where the functions can be executed in a manner not shown or discussed, including in a substantially simultaneous manner or in a reverse order according to the involved functions, which should be understood by those skilled in the art to which the embodiments of this application belong.

[0146] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0147] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0148] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0149] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0150] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A text analysis method based on a large model, characterized in that, Including: Obtain the information type corresponding to the input information; When the information type is the first information type, query in the BI analysis data table library according to the input information to obtain a first target data set. Wherein, the first information type is used to indicate that the complexity corresponding to the input information is greater than a first complexity threshold, and the similarity between each target data in the first target data set and the input information is less than a first similarity threshold and greater than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold; Perform a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set to obtain a first query data set; Input the task planning information set and the first query data set into a large model for analysis to obtain an analysis result, and generate a BI report corresponding to the input information according to the analysis result.

2. The method according to claim 1, wherein The method further includes: When the information type is the second information type, retrieve in the milvus data table library according to the input information to obtain a data table identifier set corresponding to the input information, and obtain the table structure of the data table corresponding to each data table identifier in the data table identifier set. Wherein, the second information type is used to indicate that the complexity corresponding to the input information is less than a second complexity threshold, and the second complexity threshold is less than the first complexity threshold; Use an intent recognition model to identify the input information to obtain three elements corresponding to the input information; Input the table structure of the data table corresponding to each data table identifier and the three elements corresponding to the input information into the large model to generate a Structured Query Language (SQL); Query in the data table library according to the Structured Query Language (SQL) to obtain a second query data set, and use the second query data set as the analysis result corresponding to the input information.

3. The method according to claim 1, characterized in that, The obtaining of the information type corresponding to the input information includes: When a type selection instruction of a display interface is detected, determine the information type corresponding to the input information according to the type selection instruction; Or, When the type selection instruction is not received, obtain a first vector corresponding to the input information; Use the first vector to perform multiple rounds of retrieval in the problem vector library through table routing to obtain a retrieval result set, and obtain the information type corresponding to the input information according to the retrieval result set. Wherein, the retrieval result set includes at least one retrieval result, and the similarity between each retrieval result in the at least one retrieval result and the first vector is greater than a similarity threshold.

4. The method according to claim 1, wherein The obtaining of the information type corresponding to the input information includes: When the type selection instruction is not received and the retrieval result set is not obtained, input the input information into the large model for classification processing to obtain the information type corresponding to the input information.

5. The method according to claim 1, wherein The method further includes: When the information type is the first information type, query in the BI analysis data table library according to the input information to obtain a second target data set, and use the second target data set as the analysis result corresponding to the input information, where the similarity between each target data in the second target data set and the input information is greater than a first similarity threshold.

6. The method according to claim 1, characterized in that, The step of inputting the task planning information set and the first query data set into the large model for analysis to obtain an analysis result includes: Input the task planning information set and the query data set into the large model for code generation operation to obtain data analysis code; Execute the data for data analysis to obtain an analysis result.

7. The method according to claim 1, characterized in that The step of performing a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set to obtain a first query data set includes: Obtain the application scenario information corresponding to the input information; Obtain at least one structured query language (SQL) data table library related to the application scenario; According to each task planning information in the task planning information set corresponding to the first target data set, perform a natural language query operation in each structured query language (SQL) data table in the at least one structured query language (SQL) data table library to obtain a first query data set.

8. A text analysis device based on a large model, characterized in that, It includes: A type acquisition unit for acquiring the information type corresponding to the input information; A set acquisition unit for querying in the BI analysis data table library according to the input information when the information type is the first information type to obtain a first target data set, where the first information type is used to indicate that the complexity corresponding to the input information is greater than a first complexity threshold, the similarity between each target data in the first target data set and the input information is less than the first similarity threshold and greater than a second similarity threshold, and the first similarity threshold is greater than the second similarity threshold; The set acquisition unit is further configured to perform a natural language query operation using each task planning information in the task planning information set corresponding to the first target data set to obtain a first query data set; A report generation unit for inputting the task planning information set and the first query data set into the large model for analysis to obtain an analysis result, and generating a BI report corresponding to the input information according to the analysis result.

9. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.