Data analysis processing method and device, storage medium and terminal
By using predetermined large models and virtual relationship tables for data cleaning and analysis in data analysis, multiple data table inconsistency problems are solved, improving the accuracy and efficiency of data cleaning and analysis.
Patent Information
- Application Number
- CN202510318021.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-11
AI Technical Summary
Data inconsistencies in multiple data tables lead to poor data cleaning effects, affecting the accuracy of data analysis results.
The target table is analyzed using a predetermined large model, field analysis summary information is generated, and related storage is performed in the virtual relationship table. The field analysis summary information of the table to be cleaned is imported through the virtual relationship table to be cleaned for data cleaning, and data analysis is performed based on user-defined analysis logic and natural language description.
It improves the accuracy of data cleaning effects and data analysis, simplifies the data analysis process, improves user experience, and reduces the requirements for professional skills.
Smart Images

Figure CN120296001A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data analysis, and specifically relates to a data analysis and processing method, device, storage medium, and terminal. Background Art
[0002] Data analysis is of great significance in all walks of life. Many data subjects such as enterprises or institutions often store data in multiple data tables, and the data in the multiple data tables often has problems of inconsistent data sources and formats.
[0003] These inconsistencies in the data of multiple data tables often make it impossible to effectively clean and filter the data in these multiple data tables, resulting in poor data cleaning and filtering effects, further leading to poor accuracy of the data analysis results of the data in these data tables and analysis deviations. Summary of the Invention
[0004] An embodiment of this application provides a solution that can effectively improve the data cleaning and filtering effect and improve the accuracy of data analysis.
[0005] The embodiments of this application provide the following technical solutions:
[0006] According to an embodiment of this application, a data analysis and processing method includes: analyzing user-defined analysis and processing content corresponding to a target table by using a predetermined large model to obtain field analysis summary information of at least one field, where the target table is any one of multiple data tables, and the user-defined analysis and processing content includes data analysis and processing logic; associating and storing the field analysis summary information of the at least one field with corresponding fields in a virtual relationship table, where the virtual relationship table is a table that has established field association relationships between the multiple data tables in advance; for a table to be cleaned, importing the field analysis summary information associated with the fields in the table to be cleaned from the virtual relationship table, where the table to be cleaned is any one of the multiple data tables; using the predetermined large model to perform data cleaning on the table to be cleaned according to the field analysis summary information associated with the fields in the table to be cleaned.
[0007] In some embodiments of this application, after using the predetermined large model to perform data cleaning on the table to be cleaned according to the field analysis summary information associated with the fields in the table to be cleaned, the method further includes: obtaining fields to be analyzed in a table to be analyzed according to user selection and interaction operations in a pivot table, where the table to be analyzed is any one of the multiple data tables; using the predetermined large model to generate analysis logic according to the fields to be analyzed to obtain first data analysis logic; and analyzing the data in the table to be analyzed based on the first data analysis logic to obtain a first data analysis result.
[0008] In some embodiments of the present application, after performing data cleaning on the table to be cleaned by summarizing information on fields associated with the fields in the table to be cleaned using the predetermined large model, the method further includes: obtaining natural language description content for describing data analysis requirements; importing the information on field analysis and summary associated with the fields in the natural language description content from the virtual relationship table to obtain auxiliary historical analysis information; using the predetermined large model to generate an analysis logic based on the auxiliary historical analysis information and the natural language description content to obtain a second data analysis logic; and analyzing the data in the table to be analyzed based on the second data analysis logic to obtain a second data analysis result.
[0009] In some embodiments of the present application, the method further includes: using the predetermined large model to generate a visualization chart based on the data analysis result, where the data analysis result includes the custom analysis result of the data analysis processing logic, the first data analysis result, or the second data analysis result; storing the data analysis result in a search engine, and using the predetermined large model to convert the visualization chart into a configuration file for a data visualization platform.
[0010] In some embodiments of the present application, before analyzing the user-defined analysis and processing content corresponding to the target table using the predetermined large model to obtain information on field analysis and summary of at least one field, the method further includes: analyzing the data in the target table using a data analysis processing logic custom-built by the user for the target table to obtain a custom analysis result; displaying the custom analysis result in a predetermined result display area, and obtaining the annotation information corresponding to the fields in the custom analysis result according to the annotation operations in the predetermined result display area; and determining the data analysis processing logic and the annotation information as the user-defined analysis and processing content corresponding to the target table.
[0011] In some embodiments of the present application, after importing the information on field analysis and summary associated with the fields in the table to be cleaned from the virtual relationship table and before performing data cleaning on the table to be cleaned using the information on field analysis and summary associated with the fields in the table to be cleaned by the predetermined large model, the method further includes: displaying the information on field analysis and summary in a predetermined information display area; and obtaining the edited information on field analysis and summary according to the convenient operations in the predetermined information display area.
[0012] In some embodiments of the present application, the method further includes: storing the data analysis result in a search engine, where the data analysis result includes the custom analysis result of the data analysis processing logic, the first data analysis result, or the second data analysis result; converting the data analysis result into a visual dashboard allowing secondary editing by using a large model deployed on a data visualization platform.
[0013] According to an embodiment of the present application, a data analysis processing device includes: an analysis processing module configured to: analyze user-defined analysis processing content corresponding to a target table by using a predetermined large model to obtain field analysis summary information of at least one field, where the target table is any one of multiple data tables, and the user-defined analysis processing content includes a data analysis processing logic; a deep association module configured to: associatively store the field analysis summary information of the at least one field with corresponding fields in a virtual relationship table, where the virtual relationship table is a table that has pre-established field association relationships among the multiple data tables; an information import module configured to: for a table to be cleaned, import the field analysis summary information associated with the fields in the table to be cleaned from the virtual relationship table, where the table to be cleaned is any one of the multiple data tables; a data cleaning module configured to: perform data cleaning on the table to be cleaned by using the predetermined large model according to the field analysis summary information associated with the fields in the table to be cleaned.
[0014] According to another embodiment of the present application, a storage medium stores a computer program, and when the computer program is executed by a processor of a terminal, the terminal is caused to execute the method described in the embodiments of the present application.
[0015] According to another embodiment of the present application, a terminal may include: a memory storing a computer program; a processor reading the computer program stored in the memory to execute the method described in the embodiments of the present application.
[0016] According to another embodiment of the present application, a computer program product or a computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a terminal reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the terminal executes the methods provided in various alternative implementation manners described in the embodiments of the present application.
[0017] In the embodiments of the present application, a pre - determined large - model is used to analyze the user - defined analysis and processing content corresponding to the target table, and the field analysis summary information of at least one field is obtained. The target table is any one of multiple data tables, and the user - defined analysis and processing content includes data analysis and processing logic. The field analysis summary information of the at least one field is associated and stored with the corresponding fields in a virtual relationship table. The virtual relationship table is a pre - created table that has established the field association relationships between the multiple data tables. For the table to be cleaned, the field analysis summary information associated with the fields in the table to be cleaned is imported from the virtual relationship table. The table to be cleaned is any one of the multiple data tables. The pre - determined large - model is used to perform data cleaning on the table to be cleaned according to the field analysis summary information associated with the fields in the table to be cleaned.
[0018] In this way of the embodiments of the present application, through the virtual relationship table, multiple data tables are efficiently and deeply associated with the help of a large - model. Even in the case of inconsistencies such as different sources and different formats of data in multiple data tables, reliable and effective cleaning and filtering can be achieved. The data tables after cleaning and filtering are used for data analysis, which can effectively improve the accuracy of the data analysis results. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 The flowchart of the data analysis and processing method according to an embodiment of the present application is shown.
[0021] Figure 2 The schematic diagram of the virtual relationship table according to an embodiment of the present application is shown.
[0022] Figure 3 The schematic diagram of the content analysis interface according to an embodiment of the present application is shown.
[0023] Figure 4 The schematic diagram of the data cleaning interface according to an embodiment of the present application is shown.
[0024] Figure 5 The schematic diagram of the data analysis interface according to an embodiment of the present application is shown.
[0025] Figure 6 The schematic diagram of the data analysis interface according to another embodiment of the present application is shown.
[0026] Figure 7 The block diagram of a data analysis and processing device according to an embodiment of the present application is shown.
[0027] Figure 8 The block diagram of a terminal according to an embodiment of the present application is shown. Detailed implementation manners
[0028] The present disclosure will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are only used to explain the present disclosure and are not used to limit the present disclosure. In addition, the embodiments provided below are partial embodiments for implementing the present disclosure, rather than all embodiments for implementing the present disclosure. Without conflict, the technical solutions described in the embodiments of the present disclosure can be implemented in any combination.
[0029] It should be noted that in the embodiments of the present disclosure, the term "including", "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a method or device including a series of elements not only includes the elements explicitly recited, but also includes other elements not explicitly listed, or further includes elements inherent to the implementation of the method or device. Without further limitation, an element defined by the phrase "including one..." does not exclude the existence of additional related elements in the method or device including the element (such as steps in a method or units in a device, and the units may be partial circuits, partial processors, partial programs or software, etc.).
[0030] For example, the data analysis and processing method provided in the embodiments of the present disclosure includes a series of steps, but the data analysis and processing method provided in the embodiments of the present disclosure is not limited to the recited steps. Similarly, the data analysis and processing device provided in the embodiments of the present disclosure includes a series of units, but the device provided in the embodiments of the present disclosure is not limited to including the explicitly recited units, and may further include units required for obtaining relevant information or processing based on the information.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this disclosure belongs. The terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.
[0032] It can be understood that in the specific implementation manners of the present application, when related data is involved, when the embodiments in the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards in relevant countries and regions.
[0033] Figure 1The flowchart of the data analysis and processing method according to an embodiment of the present application is schematically shown. The execution subject of the data analysis and processing method can be a terminal and / or a server with data analysis and processing capabilities. The terminal can be, for example, a computer, a calculator, a mobile phone, and a vehicle-mounted device, etc. The server can be, for example, a cloud server or a physical server, etc.
[0034] As Figure 1 shown, the data analysis and processing method may include step S110 to step S140.
[0035] Step S110, analyze the user-defined analysis and processing content corresponding to the target table by using a predetermined large model to obtain the field analysis summary information of at least one field. The target table is any one of multiple data tables, and the user-defined analysis and processing content includes data analysis and processing logic;
[0036] Step S120, associate and store the field analysis summary information of the at least one field with the corresponding field in the virtual relationship table. The virtual relationship table is a table that has established the field association relationship between the multiple data tables in advance;
[0037] Step S130, for the table to be cleaned, import the field analysis summary information associated with the fields in the table to be cleaned from the virtual relationship table. The table to be cleaned is any one of the multiple data tables;
[0038] Step S140, use the predetermined large model to perform data cleaning on the table to be cleaned according to the field analysis summary information associated with the fields in the table to be cleaned.
[0039] Create a virtual relationship table for multiple data tables. The virtual relationship table is a table that has established the field association relationship between multiple data tables. Refer to Figure 2 , for data table 210 and data table 220, a virtual relationship table 230 can be created. The virtual relationship table 230 establishes the field association relationship between data table 210 and data table 220. The virtual relationship table 230 can record the field information of the same fields in data table 210 and data table 220 (such as Figure 2 the field information of fields such as queryid and dnum shown). Among them, the virtual relationship table can adopt data tables including but not limited to the graph database Neo4j, etc.
[0040] The user configures custom data analysis and processing logic (such as SQL statements for data cleaning or data analysis) for any one target table among multiple data tables. By analyzing and processing the data in the target table with the custom data analysis and processing logic, corresponding custom analysis and processing results can be obtained, thereby obtaining the user-defined analysis and processing content corresponding to the target table. The user-defined analysis and processing content can at least include the data analysis and processing logic.
[0041] See Figure 3 , for example, the user can Figure 3 select a certain syntax (such as Spark syntax) in the selection area 310 in the interface shown in
[0042] and enter the data analysis and processing logic of the corresponding syntax (such as SparkSQL) in the display area 320. Then, by analyzing and processing the data in the target table with the data analysis and processing logic, the corresponding custom analysis and processing results can be obtained and displayed in the display area 330. Figure 3 In the display area 330 shown in
[0043] field analysis summary information of at least one field can be displayed. The field analysis summary information can at least include information such as the name, meaning, numeric type, and field application of the field. Since the virtual relationship table establishes the field association relationship between multiple data tables, the field analysis summary information of at least one field learned and summarized is further associated and stored with the corresponding fields in the virtual relationship table, and further deep association of multiple data tables can be formed. When cleaning any one of the multiple data tables that needs to be cleaned, the field analysis summary information associated with the fields in the table to be cleaned can be imported from the virtual relationship table, and a predetermined large model can be used to clean the data in the table to be cleaned according to the field analysis summary information associated with the fields in the table to be cleaned.
[0044] See Figure 4 , for example, the user can Figure 4Select "Large Language Model Processing Mode" 410 in the interface shown. At the same time, "Data Cleaning" 420 and "Table to be Cleaned" 430 can be selected. Then, the data in the table to be cleaned can be previewed in the display area 440. Further, "LLM Cleaning" can be selected to perform data cleaning on the table to be cleaned only by a predetermined large model without referring to the field analysis summary information. The cleaned data is displayed in the display area 450, and it can be seen that only one piece of data is cleaned. Further, the field analysis summary information associated with the fields in the table to be cleaned is imported and displayed in the display area 460. Then, select "LLM Cleaning + Historical Information Analysis", and use the predetermined large model to perform data cleaning on the table to be cleaned according to the field analysis summary information associated with the fields in the table to be cleaned. The cleaned data is displayed in the display area 470, and it can be seen that two pieces of data are cleaned, and the data cleaning effect is better.
[0045] In this way, through the virtual relationship table, multiple data tables are efficiently and deeply associated with the help of the large model. Even when there are inconsistency problems such as different sources and different formats in the data of multiple data tables, reliable and effective cleaning and filtering can be achieved. The data tables after cleaning and filtering are used for data analysis, which can effectively improve the accuracy of the data analysis results.
[0046] The following description Figure 1 Specific embodiments that are further optional under each step when performing data analysis and processing in the embodiment.
[0047] In one embodiment, before analyzing the user-defined analysis and processing content corresponding to the target table using the predetermined large model to obtain the field analysis summary information of at least one field, the method further includes: analyzing and processing the data in the target table using the data analysis and processing logic custom-built by the user for the target table to obtain a custom analysis and processing result; displaying the custom analysis and processing result in a predetermined result display area, and obtaining the annotation information corresponding to the fields in the custom analysis and processing result according to the annotation operation in the predetermined result display area; determining the data analysis and processing logic and the annotation information as the user-defined analysis and processing content corresponding to the target table.
[0048] Refer to Figure 3 , the user can, for example Figure 3Select a certain grammar (such as Spark grammar) in the selection area 310 of the shown interface, and customize and input the data analysis processing logic (such as SparkSQL) of the corresponding grammar in the display area 320. Then, analyze and process the data in the target table using the data analysis processing logic to obtain the corresponding custom analysis processing results and display them in the predetermined result display area 330. Further, for any field in the analysis processing results displayed in the predetermined result display area 330, the user can mark the corresponding marking information (for example, the marking information for the cilck mark can be "media asset click volume").
[0049] Jointly determine the data analysis processing logic custom-input by the user and the marking information marked by the user for the field as the custom analysis processing content corresponding to the target table. Analyze the custom analysis processing content corresponding to the target table using a predetermined large model, and more accurate field analysis summary information for at least one field can be obtained, further improving the data cleaning and filtering effect and enhancing the accuracy of data analysis.
[0050] It can be understood that in other embodiments, before analyzing the custom analysis processing content corresponding to the target table using a predetermined large model to obtain field analysis summary information for at least one field, the method further includes: analyzing and processing the data in the target table using the data analysis processing logic custom-built by the user for the target table to obtain custom analysis processing results; then, separately determining the data analysis processing logic as the custom analysis processing content corresponding to the target table.
[0051] In one embodiment, after importing the field analysis summary information associated with the fields in the to-be-cleaned table from the virtual relationship table and before performing data cleaning on the to-be-cleaned table using the predetermined large model according to the field analysis summary information associated with the fields in the to-be-cleaned table, the method further includes: displaying the field analysis summary information in a predetermined information display area; obtaining the edited field analysis summary information according to the convenient operations in the predetermined information display area.
[0052] After importing the field analysis summary information associated with the fields in the to-be-cleaned table from the virtual relationship table, display the field analysis summary information in a predetermined information display area to support the user's convenient operations (such as modifying information or adding information, etc.). According to the convenient operations in the predetermined information display area, obtain the edited field analysis summary information. For example, add the information "when the clienttype value is air_v2, the legal length of dnum is 8". Furthermore, the predetermined large model can perform data cleaning on the to-be-cleaned table according to the edited field analysis summary information, further improving the data cleaning effect.
[0053] In one embodiment, after performing data cleaning on the table to be cleaned according to the information analyzed and summarized from the fields associated with the fields in the table to be cleaned using the predetermined large model, the method may further include: obtaining the fields to be analyzed in the table to be analyzed according to the user selection interaction operations in the pivot table, where the table to be analyzed is any one of the multiple data tables; generating an analysis logic based on the fields to be analyzed using the predetermined large model to obtain a first data analysis logic; and analyzing the data in the table to be analyzed based on the first data analysis logic to obtain a first data analysis result.
[0054] Refer to Figure 5 , any one of the selected tables to be analyzed from the multiple data tables can be previewed in the display area 510. The user can perform drag-and-drop, click-and-select, and other interactive user selection operations through the pivot table. The fields to be analyzed in the table to be analyzed selected by the user selection interaction operations can be displayed in the corresponding display area 520 of the pivot table. Further, an analysis logic can be generated based on the fields to be analyzed using the predetermined large model, and the obtained first data analysis logic (data analysis SQL statement) can be displayed in the display area 530. Analyzing the data in the table to be analyzed based on the first data analysis logic can obtain a first data analysis result taking the result displayed in the display area 540 as an example.
[0055] In this implementation manner, using an intuitive pivot table method, the user can select the required fields through simple drag-and-drop and click-and-select operations, intuitively construct query conditions, and generate a first data analysis logic based on the constructed query conditions using the predetermined large model to analyze the data in the table to be analyzed, avoiding inconsistent results caused by differences in natural language descriptions. This method not only improves the user experience but also reduces the requirements for the user's professional skills, enabling more non-technical personnel to participate in data analysis.
[0056] In one embodiment, after performing data cleaning on the table to be cleaned according to the information analyzed and summarized from the fields associated with the fields in the table to be cleaned using the predetermined large model, the method may further include:
[0057] Obtaining natural language description content for describing data analysis requirements; importing the information analyzed and summarized from the fields associated with the fields in the natural language description content from the virtual relationship table to obtain auxiliary historical analysis information; generating an analysis logic based on the auxiliary historical analysis information and the natural language description content using the predetermined large model to obtain a second data analysis logic; and analyzing the data in the table to be analyzed based on the second data analysis logic to obtain a second data analysis result.
[0058] Refer to Figure 6 , the user can be in such as Figure 6Select "Large Language Model Processing Mode" 610 in the shown interface. Meanwhile, "Data Analysis" 620 and "Table to be Analyzed" 630 can be selected. Then, input the natural language description content for describing the data analysis requirements in the input area 640. The data in the table to be analyzed can be previewed in the display area 650. Further, the field analysis summary information associated with the fields in the natural language description content can be imported from the virtual relationship table, and the imported information can be displayed as auxiliary historical analysis information in the display area 660. Then, select "LLMSQL + Historical Information Analysis", and use a predetermined large model to generate analysis logic based on the auxiliary historical analysis information and the natural language description content. The generated second data analysis logic is displayed in the display area 670. Further, analyze the data in the table to be analyzed based on the second data analysis logic, and the obtained second data analysis result can be displayed in the display area 680.
[0059] In this way, for users who are not clear about the data table structure, after inputting the natural language description content for describing the data analysis requirements, by importing the auxiliary historical analysis information from the virtual data table and using a predetermined large model to generate analysis logic based on the auxiliary historical analysis information and the natural language description content, and generating the second data analysis logic for data analysis, accurate data analysis logic can be generated for data analysis even when the natural language descriptions of different users are different, further improving the data analysis accuracy of data analysis based on natural language description.
[0060] In one embodiment, after obtaining the foregoing custom analysis processing result, the first data analysis result, or the second data analysis result, the method may further include: using the predetermined large model to generate a visualization chart based on the data analysis result, where the data analysis result may include the foregoing custom analysis processing result, the first data analysis result, or the second data analysis result; storing the data analysis result in a search engine, and using the predetermined large model to convert the visualization chart into a configuration file of a data visualization platform.
[0061] Use a predetermined large model to generate a visualization chart (such as a line chart or a bar chart, etc.) based on the data analysis result, import the data analysis result into a data table in a search engine (such as ES). Then, use a predetermined large model to convert the visualization chart into a configuration file of a data visualization platform (such as Kibana), which can be stored or shared according to requirements. Other users can import this configuration file into the data visualization platform (such as Kibana) to quickly display the monitoring panel with the corresponding layout according to the original chart layout of the visualization chart. Other users do not need to reprocess the data to generate a visualization chart, further reliably improving the efficiency and response speed of visualization.
[0062] Further, in some embodiments, after generating a visualization chart based on the data analysis results using the predetermined large model, the method may further include: obtaining an adjusted chart according to the user's adjustment operation on the visualization chart.
[0063] After generating a visualization chart based on the data analysis results using a predetermined large model, multiple edits to the chart are supported. For example, the user can adjust the position, size, and appearance of the visualization chart. Information such as the type of the visualization chart can also be changed, for example, changing a line chart to a bar chart, or changing a blue line to green, etc. This further enhances the user experience.
[0064] In one embodiment, the method may further include: storing the data analysis results in a search engine, where the data analysis results include the custom analysis results of the data analysis processing logic, the first data analysis results, or the second data analysis results; using the large model deployed on the data visualization platform to convert the data analysis results into a visual dashboard that allows secondary editing.
[0065] After deploying a large model on a data visualization platform (such as Kibana) and importing the data analysis results into a search engine (such as ES), the large model deployed on the data visualization platform can directly convert the data analysis results in the search engine into a visual dashboard that allows secondary editing, without the need to be additionally converted into a configuration file.
[0066] To facilitate better implementation of the data analysis processing method provided in the embodiments of the present application, the embodiments of the present application also provide a data analysis processing device based on the above data analysis processing method. The meanings of the terms are the same as those in the above data analysis processing method, and the specific implementation details can refer to the descriptions in the method embodiments. Figure 7 The block diagram of a data analysis processing device according to an embodiment of the present application is shown.
[0067] As Figure 7As shown in the figure, the data analysis and processing device 700 may include: The analysis and processing module 710 may be configured to: analyze the user-defined analysis and processing content corresponding to the target table by using a predetermined large model to obtain the field analysis summary information of at least one field, where the target table is any one of multiple data tables, and the user-defined analysis and processing content includes data analysis and processing logic; The deep association module 720 may be configured to: associate and store the field analysis summary information of the at least one field with the corresponding fields in the virtual relationship table, where the virtual relationship table is a pre-created table that establishes the field association relationship between the multiple data tables; The information import module 730 may be configured to: for the table to be cleaned, import the field analysis summary information associated with the fields in the table to be cleaned from the virtual relationship table, where the table to be cleaned is any one of the multiple data tables; The data cleaning module 740 may be configured to: perform data cleaning on the table to be cleaned by using the predetermined large model according to the field analysis summary information associated with the fields in the table to be cleaned.
[0068] In some embodiments of the present application, after performing data cleaning on the table to be cleaned by using the predetermined large model according to the field analysis summary information associated with the fields in the table to be cleaned, the device further includes a first logic analysis module, configured to: obtain the fields to be analyzed in the table to be analyzed according to the user selection and interaction operations in the pivot table, where the table to be analyzed is any one of the multiple data tables; generate analysis logic by using the predetermined large model according to the fields to be analyzed to obtain the first data analysis logic; analyze the data in the table to be analyzed based on the first data analysis logic to obtain the first data analysis result.
[0069] In some embodiments of the present application, after performing data cleaning on the table to be cleaned by using the predetermined large model according to the field analysis summary information associated with the fields in the table to be cleaned, the device further includes a second logic analysis module, configured to: obtain the natural language description content for describing the data analysis requirements; import the field analysis summary information associated with the fields in the natural language description content from the virtual relationship table to obtain the auxiliary historical analysis information; generate analysis logic by using the predetermined large model according to the auxiliary historical analysis information and the natural language description content to obtain the second data analysis logic; analyze the data in the table to be analyzed based on the second data analysis logic to obtain the second data analysis result.
[0070] In some embodiments of the present application, the device further includes a visualization processing module, configured to: generate a visualization chart based on the data analysis result by using the predetermined large model, where the data analysis result includes the custom analysis result of the data analysis processing logic, the first data analysis result, or the second data analysis result; store the data analysis result in a search engine, and convert the visualization chart into a configuration file of the data visualization platform by using the predetermined large model.
[0071] In some embodiments of the present application, before analyzing the user-defined analysis processing content corresponding to the target table by using the predetermined large model to obtain the field analysis summary information of at least one field, the device further includes an information acquisition module, configured to: analyze and process the data in the target table by using the data analysis processing logic custom-built by the user for the target table to obtain a custom analysis result; display the custom analysis result in a predetermined result display area, and obtain the annotation information corresponding to the fields in the custom analysis result according to the annotation operation in the predetermined result display area; determine the data analysis processing logic and the annotation information as the user-defined analysis processing content corresponding to the target table.
[0072] In some embodiments of the present application, after importing the field analysis summary information associated with the fields in the to-be-cleaned table from the virtual relationship table, before performing data cleaning on the to-be-cleaned table by using the predetermined large model according to the field analysis summary information associated with the fields in the to-be-cleaned table, the device further includes an information editing module, configured to: display the field analysis summary information in a predetermined information display area; obtain the edited field analysis summary information according to the convenient operations in the predetermined information display area.
[0073] In some embodiments of the present application, the device further includes a post-processing module, configured to: store the data analysis result in a search engine, where the data analysis result includes the custom analysis result of the data analysis processing logic, the first data analysis result, or the second data analysis result; convert the data analysis result into a visual dashboard allowing secondary editing by using the large model carried on the data visualization platform.
[0074] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0075] In addition, an embodiment of the present application further provides a terminal, such as Figure 8 shown Figure 8 which shows a block diagram of a terminal according to an embodiment of the present application. Specifically:
[0076] The terminal may include a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, a power supply 803, an input unit 804, and other components. Those skilled in the art can understand that Figure 8 the terminal structure shown in
[0077] does not limit the terminal, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Among them:
[0078] The processor 801 is the control center of the terminal, connecting various parts of the entire computer device through various interfaces and lines. By running or executing software programs and / or modules stored in the memory 802, and calling data stored in the memory 802, it executes various functions of the computer device and processes data, thereby monitoring the terminal as a whole. Optionally, the processor 801 may include one or more processing cores; preferably, the processor 801 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 801.
[0079] The terminal further includes a power supply 803 that powers each component. Preferably, the power supply 803 can be logically connected to the processor 801 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 803 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0080] The terminal may further include an input unit 804, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0081] Although not shown, the terminal may further include a display unit and the like, which will not be elaborated here. Specifically, in this embodiment, the processor 801 in the terminal will load the executable files corresponding to the processes of one or more computer programs into the memory 802 according to the following instructions, and the processor 801 will run the computer programs stored in the memory 802 to implement various functions in the foregoing embodiments of the present application. For example, the processor 801 can execute the following steps:
[0082] Analyze the user-defined analysis and processing content corresponding to the target table using a predetermined large model to obtain field analysis summary information for at least one field. The target table is any one of multiple data tables, and the user-defined analysis and processing content includes data analysis and processing logic; store the field analysis summary information for the at least one field in association with the corresponding fields in a virtual relationship table, where the virtual relationship table is a pre-created table that establishes the field association relationships between the multiple data tables; for the table to be cleaned, import the field analysis summary information associated with the fields in the table to be cleaned from the virtual relationship table, where the table to be cleaned is any one of the multiple data tables; use the predetermined large model to perform data cleaning on the table to be cleaned based on the field analysis summary information associated with the fields in the table to be cleaned.
[0083] Those of ordinary skill in the art can understand that all or part of the steps in the above-mentioned various methods can be completed by a computer program, or by controlling relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0084] Therefore, the embodiments of the present application further provide a storage medium, in which a computer program is stored, and the computer program can be loaded by a processor to execute the steps in any one of the methods provided by the embodiments of the present application.
[0085] Among them, the storage medium may be a computer-readable storage medium, and the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0086] Since the computer program stored in the storage medium can execute the steps in any of the methods provided in the embodiments of the present application, the beneficial effects achievable by the methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.
[0087] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the embodiments disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.
[0088] It should be understood that the present application is not limited to the embodiments already described and shown in the drawings, but various modifications and changes can be made without departing from its scope.
Claims
1. A data analysis and processing method, characterized in that, Including: Analyze the user-defined analysis and processing content corresponding to the target table using a predetermined large model to obtain field analysis summary information of at least one field. The target table is any one of multiple data tables, and the user-defined analysis and processing content includes data analysis and processing logic. Associate and store the field analysis summary information of the at least one field with the corresponding fields in a virtual relationship table. The virtual relationship table is a pre-created table that establishes the field association relationships between the multiple data tables. For the table to be cleaned, import the field analysis summary information associated with the fields in the table to be cleaned from the virtual relationship table. The table to be cleaned is any one of the multiple data tables. Use the predetermined large model to perform data cleaning on the table to be cleaned according to the field analysis summary information associated with the fields in the table to be cleaned.
2. The method according to claim 1, wherein After using the predetermined large model to perform data cleaning on the table to be cleaned according to the field analysis summary information associated with the fields in the table to be cleaned, the method further includes: Obtain the fields to be analyzed in the table to be analyzed according to the user selection and interaction operations in the pivot table. The table to be analyzed is any one of the multiple data tables. Use the predetermined large model to generate analysis logic based on the fields to be analyzed to obtain the first data analysis logic. Analyze the data in the table to be analyzed based on the first data analysis logic to obtain the first data analysis result.
3. The method according to claim 1, characterized in that, After using the predetermined large model to perform data cleaning on the table to be cleaned according to the field analysis summary information associated with the fields in the table to be cleaned, the method further includes: Obtain the natural language description content for describing the data analysis requirements. Import the field analysis summary information associated with the fields in the natural language description content from the virtual relationship table to obtain auxiliary historical analysis information. Use the predetermined large model to generate analysis logic based on the auxiliary historical analysis information and the natural language description content to obtain the second data analysis logic. Analyze the data in the table to be analyzed based on the second data analysis logic to obtain the second data analysis result.
4. The method according to claim 2 or 3, characterized in that, The method further includes: Use the predetermined large model to generate a visualization chart based on the data analysis result. The data analysis result includes the custom analysis and processing result of the data analysis and processing logic, the first data analysis result, or the second data analysis result. Store the data analysis result in a search engine and use the predetermined large model to convert the visualization chart into a configuration file for a data visualization platform.
5. The method according to claim 1, wherein Before using the predetermined large model to analyze the user-defined analysis and processing content corresponding to the target table to obtain the field analysis summary information of at least one field, the method further includes: Use the data analysis and processing logic custom-built by the user for the target table to analyze and process the data in the target table to obtain a custom analysis and processing result. Display the custom analysis and processing result in a predetermined result display area and obtain the annotation information corresponding to the fields in the custom analysis and processing result according to the annotation operations in the predetermined result display area. Determine the data analysis and processing logic and the annotation information as the user-defined analysis and processing content corresponding to the target table.
6. The method according to claim 1, wherein After importing the field analysis and summary information associated with the fields in the to-be-cleaned table from the virtual relationship table, and before using the predetermined large model to perform data cleaning on the to-be-cleaned table according to the field analysis and summary information associated with the fields in the to-be-cleaned table, the method further includes: Display the field analysis and summary information in a predetermined information display area; Obtain the edited field analysis and summary information according to the convenient operations in the predetermined information display area.
7. The method according to claim 2 or 3, characterized in that, The method further includes: Store the data analysis result in a search engine, where the data analysis result includes the custom analysis and processing result of the data analysis and processing logic, the first data analysis result, or the second data analysis result; Use the large model carried on the data visualization platform to convert the data analysis result into a visual dashboard that allows secondary editing.
8. A data analysis and processing device, characterized in that, Include: An analysis and processing module, configured to: use a predetermined large model to analyze the user-defined analysis and processing content corresponding to the target table to obtain the field analysis and summary information of at least one field, where the target table is any one of multiple data tables, and the user-defined analysis and processing content includes data analysis and processing logic; A deep association module, configured to: associate and store the field analysis and summary information of the at least one field with the corresponding fields in a virtual relationship table, where the virtual relationship table is a pre-created table that establishes the field association relationship between the multiple data tables; An information import module, configured to: for the to-be-cleaned table, import the field analysis and summary information associated with the fields in the to-be-cleaned table from the virtual relationship table, where the to-be-cleaned table is any one of the multiple data tables; A data cleaning module, configured to: use the predetermined large model to perform data cleaning on the to-be-cleaned table according to the field analysis and summary information associated with the fields in the to-be-cleaned table.
9. A storage medium, characterized in that, It stores a computer program, and when the computer program is executed by the processor of the terminal, the terminal executes the method according to any one of claims 1 to 7.
10. A terminal, characterized in that, Include: A memory, storing a computer program; A processor, reading the computer program stored in the memory to execute the method according to any one of claims 1 to 7.