Data analysis method, system and equipment based on large language model and storage medium

Through a data analysis method based on a large language model, natural language text is processed to generate domain-specific language texts, and the problems of low data analysis efficiency and wrong results in the prior art are solved, and efficient and accurate data analysis is achieved.

CN119917804APending Publication Date: 2025-05-02SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510009252.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing data analysis processing methods are highly dependent on manual labor, which leads to low efficiency of data analysis and errors in analysis results caused by manual errors.

Method used

The data analysis method based on the large language model is adopted to determine the data analysis task by receiving the natural language text sent by the client, and process the natural language text based on the preset large language model to obtain the domain-specific language text, and perform corresponding processing in the query task or attribution task to generate the analysis results.

Benefits of technology

It improves the efficiency of data analysis and the accuracy of results, reduces the dependence on labor, and reduces the possibility of errors in analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917804A_ABST
    Figure CN119917804A_ABST
Patent Text Reader

Abstract

The invention provides a data analysis method, system and device based on a large language model, and a storage medium. The method comprises the following steps: receiving a natural language text sent by a client; determining a data analysis task corresponding to the natural language text; processing the natural language text based on a preset large language model to obtain a domain-specific language text; under the condition that the data analysis task is a query task, based on a preset database, obtaining a query result corresponding to the domain specific language text; under the condition that the data analysis task is an attribution task, attribution processing is conducted on the domain specific language text, and an attribution result is obtained; and sending the query result or the attribution result to the client. In the data analysis and processing process, the natural language text input by the user is processed through the large language model and different processing links without depending on manpower, so that the data analysis efficiency and the result accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of natural language processing technology, and in particular to a data analysis method, system, device and storage medium based on a large language model. Background Art

[0002] Currently, in the field of data analysis, users can send analysis requests and text to the server through the client. The data analyst on the server will collect corresponding data based on the text sent by the user, generate analysis results, and feed back the above analysis results to the client.

[0003] The above data analysis and processing method is highly dependent on manual labor, which leads to low efficiency of data analysis and the possibility of erroneous analysis results due to manual errors. Summary of the Invention

[0004] The main purpose of this application is to provide a data analysis method, system, device and storage medium, aiming to solve the technical problems that the existing data analysis processing method is highly dependent on manual labor, resulting in low data analysis efficiency and erroneous analysis results due to human errors.

[0005] To achieve the above objectives, the present application provides a data analysis method, which includes:

[0006] Receive natural language text sent by the client;

[0007] Determining a data analysis task corresponding to the natural language text;

[0008] Processing the natural language text based on a preset large language model to obtain a domain-specific language text;

[0009] In the case where the data analysis task is a query task, obtaining a query result corresponding to the domain-specific language text based on a preset database;

[0010] In a case where the data analysis task is an attribution task, performing attribution processing on the domain-specific language text to obtain an attribution result;

[0011] The query result or the attribution result is sent to the client.

[0012] Optionally, determining the data analysis task corresponding to the natural language text includes:

[0013] classifying the natural language text;

[0014] Determining a data analysis task corresponding to the natural language text according to the category to which the natural language text belongs;

[0015] The data analysis tasks include query tasks and attribution tasks.

[0016] Optionally, the processing of the natural language text based on a preset large language model to obtain a domain-specific language text includes:

[0017] Determining a prompt word corresponding to the natural language text according to a data analysis task corresponding to the natural language text;

[0018] The natural language text and the prompt words corresponding to the natural language text are input into the large language model to obtain domain-specific language text.

[0019] Optionally, obtaining the query result corresponding to the domain-specific language text based on a preset database includes:

[0020] Correcting the domain-specific language text based on a preset metadata database;

[0021] converting the corrected domain-specific language text into structured query language text;

[0022] The structured query language text is queried in the preset database to obtain query results corresponding to the domain-specific language text.

[0023] Optionally, performing attribution processing on the domain-specific language text to obtain an attribution result includes:

[0024] Obtaining multiple data dimensions corresponding to the domain-specific language text;

[0025] The data corresponding to the multiple data dimensions are obtained through a preset attribution server, and an attribution result is generated based on the data corresponding to the multiple data dimensions.

[0026] Optionally, before sending the query result or the attribution result to the client, the method further includes:

[0027] Receive modification instructions;

[0028] In response to the modification instruction, the query result or the attribution result is modified.

[0029] Optionally, sending the query result or the attribution result to the client includes:

[0030] Inputting the natural language text into the large language model to obtain analysis intent;

[0031] Sending the analysis intention to the client;

[0032] Upon receiving a confirmation instruction sent by the client, sending the query result or the attribution result to the client;

[0033] The confirmation instruction is used to indicate that the analysis intention is correct.

[0034] In addition, to achieve the above objectives, the present application also provides a data analysis system, which includes an analysis server, a large language model, and an attribution server. The analysis server is respectively connected to the large language model and the attribution server, and the analysis server is also connected to an external client:

[0035] The analysis server is configured to receive the natural language text sent by the client;

[0036] Determining a data analysis task corresponding to the natural language text;

[0037] The large language model is used to process the natural language text to obtain domain-specific language text;

[0038] The analysis server is further configured to obtain, based on a preset database, a query result corresponding to the domain-specific language text when the data analysis task is a query task;

[0039] The attribution server is configured to perform attribution processing on the domain-specific language text to obtain an attribution result when the data analysis task is an attribution task;

[0040] The analysis server is further configured to send the query result or the attribution result to the client.

[0041] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0042] The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the data analysis methods proposed in the embodiments of the present application are implemented.

[0043] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0044] The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of any one of the data analysis methods proposed in the embodiments of the present application.

[0045] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0046] The present application provides a data analysis method, system, device and storage medium based on a large language model, the method comprising: receiving a natural language text sent by a client; determining a data analysis task corresponding to the natural language text; processing the natural language text based on a preset large language model to obtain a domain-specific language text; in the case where the data analysis task is a query task, obtaining a query result corresponding to the domain-specific language text based on a preset database; in the case where the data analysis task is an attribution task, performing attribution processing on the domain-specific language text to obtain an attribution result; and sending the query result or attribution result to the client. In an embodiment of the present application, the natural language text is processed according to a preset large language model to obtain a domain-specific language text, and then, based on the type of data analysis task, different domain-specific language texts are processed through different processing links to obtain data analysis results; the above-mentioned data analysis processing process processes the natural language text input by the user through a large language model and different processing links, without relying on manual labor, thereby improving the efficiency of data analysis and the accuracy of the results. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0048] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0049] Figure 2 is a flow chart of the data analysis method provided in an embodiment of the present application;

[0050] Figure 3 is an application flow chart of the data analysis method provided in an embodiment of the present application;

[0051] Figure 4 This is a schematic diagram of the structure of an embodiment of the data analysis system provided in the embodiments of the present application;

[0052] Figure 5 This is a basic structural block diagram of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0053] The data analysis method provided in the embodiments of the present application is applied to a data analysis system. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by technicians in the technical field of this application; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" in the specification and claims of this application and the above-mentioned figure descriptions and any variations thereof are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned figures are used to distinguish different objects, not to describe a specific order.

[0054] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0055] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0056] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0057] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social online platform software, etc.

[0058] Terminal devices 101, 102, and 103 can be various electronic devices with display screens and support web browsing, including but not limited to smartphones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV), laptop computers, desktop computers, etc.

[0059] The server 105 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal devices 101 , 102 , and 103 .

[0060] It should be noted that the data analysis method provided in the embodiments of the present application is generally executed by a server / terminal device, and accordingly, the data analysis system is generally set in the server / terminal device.

[0061] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0062] Please refer to Figure 2 , shows a flow chart of an embodiment of the data analysis method proposed in this application. The embodiment of this application can acquire and process relevant data based on artificial intelligence technology.

[0063] It should be noted that the data analysis method provided in the embodiment of the present application can be applied to a data analysis system. The above-mentioned data analysis system includes an analysis server, a large language model and an attribution server. The analysis server is respectively communicated with the large language model and the attribution server, and the analysis server is also communicated with an external client.

[0064] The data analysis method provided in the embodiment of the present application includes the following steps:

[0065] S210: Receive natural language text sent by the client.

[0066] In this step, the user can perform corresponding operations on the client to generate natural language text and send the natural language text to the analysis server.

[0067] Specifically, the client displays an input interface, and the user enters natural language text in the input interface. After receiving the natural language text, the client sends the natural language text to the analysis server.

[0068] In this step, a natural language input framework is provided. Users only need to input natural language to perform data analysis without having to master complex data query languages, thereby lowering the technical threshold.

[0069] S220: Determine a data analysis task corresponding to the natural language text.

[0070] In this step, after receiving the natural language text, the data analysis task corresponding to the natural language text is determined. For specific implementation methods, please refer to the subsequent embodiments.

[0071] Optionally, the above-mentioned data analysis tasks include but are not limited to query tasks and analysis tasks, wherein the above-mentioned query task representation is based on natural language text for data query, and the above-mentioned analysis task representation is based on natural language text for data statistics and analysis in multiple dimensions.

[0072] S230: Process the natural language text based on a preset large language model to obtain a domain-specific language text.

[0073] In this step, a trained Large Language Model (LLM) model is pre-set, and the natural language text is input into the LLM to obtain the domain-specific language text.

[0074] S240 , when the data analysis task is a query task, obtaining a query result corresponding to the domain-specific language text based on a preset database.

[0075] One possible scenario is that the data analysis task corresponding to the natural language text is a query task. In this case, a query is performed on the domain-specific language text in a preset database to obtain a query result. Specific implementation methods are described in the subsequent examples.

[0076] S250: When the data analysis task is an attribution task, attribution processing is performed on the domain-specific language text to obtain an attribution result.

[0077] Another possible scenario is that the data analysis task corresponding to the natural language text is an attribution task. In this case, the attribution server is called to perform attribution processing on the domain-specific language text to obtain the attribution result. For specific implementation methods, please refer to the subsequent examples.

[0078] S260: Send the query result or the attribution result to the client.

[0079] In this step, after obtaining the query result or attribution result, the query result or attribution result is sent to the client.

[0080] In an embodiment of the present application, natural language text is processed according to a preset large language model to obtain domain-specific language text, and then based on the type of data analysis task, different domain-specific language texts are processed through different processing links to obtain data analysis results; the above-mentioned data analysis and processing process processes the natural language text input by the user through a large language model and different processing links, without relying on manual labor, thereby improving the efficiency of data analysis and the accuracy of the results.

[0081] Optionally, determining the data analysis task corresponding to the natural language text includes:

[0082] classifying the natural language text;

[0083] Determining a data analysis task corresponding to the natural language text according to the category to which the natural language text belongs;

[0084] The data analysis tasks include query tasks and attribution tasks.

[0085] In this embodiment, after the natural language text is obtained, the natural language text is classified.

[0086] Alternatively, vector machines, maximum entropy models, or naive Bayes algorithms can be used to classify natural language text.

[0087] In this embodiment, data analysis tasks include two categories, namely query tasks and attribution tasks. According to the category to which the natural language text belongs, the data analysis task corresponding to the natural language text is determined.

[0088] For example, the natural language text is "What is the total monthly subsidy amount for the top 10 cities with the highest number of matching orders in September?" The data analysis task corresponding to this natural language text is a query task.

[0089] For another example, the natural language text is "Why did the order volume fluctuate on September 2nd?" The data analysis task corresponding to this natural language text is an attribution task.

[0090] In this embodiment, the natural language text provided by the user is classified, the data analysis task corresponding to the natural language text is determined, and then the data analysis task is executed through an appropriate processing link, thereby improving the efficiency of data analysis.

[0091] Optionally, the processing of the natural language text based on a preset large language model to obtain a domain-specific language text includes:

[0092] Determining a prompt word corresponding to the natural language text according to a data analysis task corresponding to the natural language text;

[0093] The natural language text and the prompt words corresponding to the natural language text are input into the large language model to obtain domain-specific language text.

[0094] In this embodiment, different prompt words are set for different categories of natural language text. The prompt words can be used to call the processing link corresponding to the natural language processing text to realize the processing of different categories of data analysis tasks and improve the efficiency of data processing.

[0095] In this embodiment, the natural language text and the prompt words corresponding to the natural language text are input into the large language model to obtain the domain-specific language text.

[0096] Optionally, obtaining the query result corresponding to the domain-specific language text based on a preset database includes:

[0097] Correcting the domain-specific language text based on a preset metadata database;

[0098] converting the corrected domain-specific language text into structured query language text;

[0099] The structured query language text is queried in the preset database to obtain query results corresponding to the domain-specific language text.

[0100] In this embodiment, a metadata database is pre-set, and after obtaining the domain-specific language text, the domain-specific language text is corrected based on the database. An optional implementation method is to use the metadata database to perform synonym correction on the domain-specific language text. Specifically, a field in a domain-specific language text is selected as the target of synonym correction, and a controlled vocabulary is selected. By adding a synonym service, the vocabulary is identified and associated. When starting synonym correction, a panel for judging the results is created, which displays the number of successfully matched rows and the score of the best candidate. Finally, synonyms after synonym correction are obtained by editing columns and adding column-based GREL expressions.

[0101] Another alternative implementation is to implement domain-specific language text correction using a meta-index dictionary in the metadata repository. Specifically, the meta-index dictionary normalizes the input text using a sub-dictionary to check for phrase matches. If a sub-dictionary fails to recognize a word, an error is reported.

[0102] In this embodiment, the domain-specific language text is corrected through the metadata database to obtain accurate domain-specific language text, thereby improving the accuracy of the query result.

[0103] In existing technologies, after obtaining natural language text, it is typically directly converted into Structured Query Language (SQL) text, which is then used to query the database to obtain results. However, the accuracy of converting natural language text into SQL text is low, reaching only 60% at best. This leads to errors in the final data analysis results.

[0104] In order to solve the above-mentioned technical problems, in this embodiment, after obtaining the domain-specific language text, the domain-specific language text is corrected, and the corrected domain-specific language text is converted into a structured query language text, and then the structured query language text is queried in a preset database; since the accuracy of converting from the domain-specific language text to the structured query language text is high, the errors in the query results are reduced and the accuracy of the query results is improved.

[0105] In addition, some query results may contain a lot of redundant information, such as the order volume of each city. If only certain cities have business development, the results may contain a lot of null values ​​or data with no business value. In this embodiment, redundant data will be intercepted according to the rules of the data analysis business.

[0106] Optionally, performing attribution processing on the domain-specific language text to obtain an attribution result includes:

[0107] Obtaining multiple data dimensions corresponding to the domain-specific language text;

[0108] The data corresponding to the multiple data dimensions are obtained through a preset attribution server, and an attribution result is generated based on the data corresponding to the multiple data dimensions.

[0109] It should be noted that a domain-specific language text corresponds to multiple data dimensions. For example, a domain-specific language text may correspond to the data dimensions of time, order quantity, and order amount.

[0110] In this embodiment, by parsing the domain-specific language text, multiple data dimensions corresponding to the domain-specific language text are obtained, and the attribution server is called to obtain data corresponding to the multiple data dimensions. For example, for the data dimension representing time, the attribution server is called to obtain multiple data on the time dimension.

[0111] Furthermore, the attribution server can generate a report based on the data corresponding to the multiple data dimensions through data increase and decrease logic, and determine the report as the attribution result. The specific implementation method is not elaborated here.

[0112] In this embodiment, in the data attribution scenario, data of multiple data dimensions corresponding to domain-specific language texts are applied to generate reports through data addition and subtraction logic, providing in-depth data insights.

[0113] Optionally, before sending the query result or the attribution result to the client, the method further includes:

[0114] Receive modification instructions;

[0115] In response to the modification instruction, the query result or the attribution result is modified.

[0116] In this embodiment, after generating the query results or attribution results, the data analyst views the query results or attribution results through the analysis server. The data analyst can perform corresponding operations on the analysis server to generate modification instructions. The analysis server modifies the query results or attribution results based on the modification instructions.

[0117] In this embodiment, after generating the query results or attribution results, the data analyst can also manually polish and verify the query results or attribution results by sending modification instructions to the analysis server to avoid erroneous data and ensure the accuracy of the results.

[0118] Optionally, sending the query result or the attribution result to the client includes:

[0119] Inputting the natural language text into the large language model to obtain analysis intent;

[0120] Sending the analysis intention to the client;

[0121] Upon receiving a confirmation instruction sent by the client, sending the query result or the attribution result to the client;

[0122] The confirmation instruction is used to indicate that the analysis intention is correct.

[0123] In this embodiment, after receiving the natural language text, the natural language text is also input into the large language model to obtain the analysis intent, and the analysis intent is sent to the client. The large language model can be used to perform word meaning analysis on the natural language text to obtain the analysis intent.

[0124] After receiving the analysis intent, the client can display the analysis intent in text form and receive instructions generated by the client based on the analysis intent.

[0125] Specifically, after receiving the analysis intent, the client displays the analysis intent and two controllable components, one of which indicates a correct analysis intent and the other indicates an incorrect analysis intent. If the user clicks the component indicating a correct analysis intent, a confirmation instruction is sent to the analysis server, which then sends the query results or attribution results to the client. If the user clicks the component indicating an incorrect analysis intent, an error instruction is sent to the analysis server, which then sends a prompt message to the client, asking the user to re-enter the natural language text.

[0126] To understand the overall technical solution, please refer to Figure 3 , Figure 3 The user, analysis service backend, big model and attribution service backend are shown, wherein the above-mentioned user is the client in the above-mentioned embodiment, the above-mentioned big model is the big language model in the above-mentioned embodiment, the above-mentioned analysis service backend is also called the analysis server, and the above-mentioned attribution service backend is also called the attribution server.

[0127] in, Figure 3 The prompt in the example is the prompt word in the embodiment of the present application. Figure 3 The DSL in the embodiment of the present application is the domain specific language text, Figure 3 The SQL in the embodiment of the present application is the structured query language text.

[0128] like Figure 3 As shown, the client sends natural language text to the analysis server, and the analysis server inputs the natural language text and the prompt words corresponding to the natural language text into the large language model to obtain domain-specific language text.

[0129] When the analysis task corresponding to the natural language text is a query task, the analysis server verifies and supplements the domain-specific language text through the metadata database, converts the domain-specific language text into a structured query language text, performs a query in the database, and obtains the query results.

[0130] When the analysis task corresponding to the natural language text is an analysis task, the analysis server extracts each data dimension in the domain-specific language text and sends the data on each data dimension to the attribution server. The attribution server sends the generated attribution report, that is, the attribution result, to the analysis server.

[0131] After obtaining the query results or attribution results, the analysis server inputs the natural language text and the prompt words corresponding to the natural language text into the large language model to obtain the analysis intent, and sends the analysis intent and data analysis results to the client.

[0132] See also Figure 4The present application provides a data analysis system 400, which includes an analysis server 410, a large language model 420, and an attribution server 430. The analysis server 410 is respectively connected to the large language model 420 and the attribution server 430. The analysis server 410 is also connected to an external client:

[0133] The analysis server 410 is configured to receive the natural language text sent by the client;

[0134] Determining a data analysis task corresponding to the natural language text;

[0135] The large language model 420 is used to process the natural language text to obtain domain-specific language text;

[0136] The analysis server 410 is further configured to obtain a query result corresponding to the domain-specific language text based on a preset database when the data analysis task is a query task;

[0137] The attribution server 430 is configured to perform attribution processing on the domain-specific language text to obtain an attribution result when the data analysis task is an attribution task;

[0138] The analysis server 410 is further configured to send the query result or the attribution result to the client.

[0139] Optionally, the analysis server 410 is specifically configured to:

[0140] classifying the natural language text;

[0141] Determining a data analysis task corresponding to the natural language text according to the category to which the natural language text belongs;

[0142] The data analysis tasks include query tasks and attribution tasks.

[0143] Optionally, the large language model 420 is specifically used to:

[0144] Determining a prompt word corresponding to the natural language text according to a data analysis task corresponding to the natural language text;

[0145] The natural language text and the prompt words corresponding to the natural language text are input into the large language model to obtain domain-specific language text.

[0146] Optionally, the analysis server 410 is specifically configured to:

[0147] Correcting the domain-specific language text based on a preset metadata database;

[0148] converting the corrected domain-specific language text into structured query language text;

[0149] The structured query language text is queried in the preset database to obtain query results corresponding to the domain-specific language text.

[0150] Optionally, the attribution server 430 is specifically configured to:

[0151] Obtaining multiple data dimensions corresponding to the domain-specific language text;

[0152] The data corresponding to the multiple data dimensions are obtained through a preset attribution server, and an attribution result is generated based on the data corresponding to the multiple data dimensions.

[0153] Optionally, the analysis server 410 is further configured to:

[0154] Receive modification instructions;

[0155] In response to the modification instruction, the query result or the attribution result is modified.

[0156] Optionally, the analysis server 410 is specifically configured to:

[0157] Inputting the natural language text into the large language model to obtain analysis intent;

[0158] Sending the analysis intention to the client;

[0159] Upon receiving a confirmation instruction sent by the client, sending the query result or the attribution result to the client;

[0160] The confirmation instruction is used to indicate that the analysis intention is correct.

[0161] To solve the above technical problems, the present application also provides a computer device. Figure 5 , Figure 5 This is a basic structural block diagram of the computer device in this embodiment.

[0162] The computer device 5 includes a memory 51, a processor 52, and a network interface 53 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 5 with components 51-53, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0163] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0164] The memory 51 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, magnetic disk, optical disk, etc. In some embodiments, the memory 51 can be an internal storage unit of the computer device 5, such as the hard disk or memory of the computer device 5. In other embodiments, the memory 51 can also be an external storage device of the computer device 5, such as a plug-in hard disk equipped on the computer device 5, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc. Of course, the memory 51 can also include both the internal storage unit of the computer device 5 and its external storage device. In this embodiment, the memory 51 is generally used to store the operating system and various application software installed on the computer device 5, such as the program code of the data analysis method. In addition, the memory 51 can also be used to temporarily store various types of data that have been output or are to be output.

[0165] In some embodiments, the processor 52 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 52 is generally used to control the overall operation of the computer device 5. In this embodiment, the processor 52 is used to execute program code stored in the memory 51 or process data, such as executing the program code of the data analysis method.

[0166] The network interface 53 may include a wireless network interface or a wired network interface. The network interface 53 is generally used to establish a communication connection between the computer device 5 and other electronic devices.

[0167] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores the data analysis program, and the data analysis program can be executed by at least one processor to enable the at least one processor to perform the steps of the data analysis method as described above.

[0168] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware online platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0169] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0170] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions described in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the present application specification and the accompanying drawings, directly or indirectly used in other related technical fields, is also within the scope of patent protection of the present application.

Claims

1. A data analysis method based on a large language model, characterized in that: The method comprises: Receive natural language text sent by the client; Determining a data analysis task corresponding to the natural language text; Processing the natural language text based on a preset large language model to obtain a domain-specific language text; In the case where the data analysis task is a query task, obtaining a query result corresponding to the domain-specific language text based on a preset database; In the case where the data analysis task is an attribution task, performing attribution processing on the domain-specific language text to obtain an attribution result; The query result or the attribution result is sent to the client.

2. The method according to claim 1, characterized in that The determining of the data analysis task corresponding to the natural language text includes: Classifying the natural language text; Determining a data analysis task corresponding to the natural language text according to the category to which the natural language text belongs; The data analysis tasks include query tasks and attribution tasks.

3. The method according to claim 1, characterized in that The processing of the natural language text based on the preset large language model to obtain the domain-specific language text includes: Determining a prompt word corresponding to the natural language text according to a data analysis task corresponding to the natural language text; The natural language text and the prompt words corresponding to the natural language text are input into the large language model to obtain a domain-specific language text.

4. The method according to claim 1, characterized in that: The obtaining of the query result corresponding to the domain-specific language text based on the preset database includes: Correcting the domain-specific language text based on a preset metadata database; converting the corrected domain specific language text into structured query language text; The structured query language text is queried in the preset database to obtain query results corresponding to the field-specific language text.

5. The method according to claim 1, characterized in that The attribution processing is performed on the domain-specific language text to obtain an attribution result, including: Obtaining multiple data dimensions corresponding to the domain-specific language text; The data corresponding to the multiple data dimensions are obtained through a preset attribution server, and attribution results are generated based on the data corresponding to the multiple data dimensions.

6. The method according to claim 1, characterized in that Before sending the query result or the attribution result to the client, the method further includes: Receive modification instructions; In response to the modification instruction, the query result or the attribution result is modified.

7. The method according to any one of claims 1 to 6, characterized in that The sending the query result or the attribution result to the client includes: Inputting the natural language text into the large language model to obtain analysis intent; Sending the analysis intention to the client; Upon receiving a confirmation instruction sent by the client, sending the query result or the attribution result to the client; The confirmation instruction is used to indicate that the analysis intention is correct.

8. A data analysis system, characterized in that: The system includes an analysis server, a large language model and an attribution server. The analysis server is connected to the large language model and the attribution server respectively, and the analysis server is also connected to an external client: The analysis server is used to receive the natural language text sent by the client; Determining a data analysis task corresponding to the natural language text; The large language model is used to process the natural language text to obtain a domain-specific language text; The analysis server is further configured to obtain, when the data analysis task is a query task, a query result corresponding to the domain-specific language text based on a preset database; The attribution server is used to perform attribution processing on the domain-specific language text to obtain an attribution result when the data analysis task is an attribution task; The analysis server is also used to send the query result or the attribution result to the client.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the data analysis method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data analysis method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Data analysis method and system based on large model

    CN121388128A