Data analysis method, device, storage medium and electronic device
By configuring controls for data relationships in the canvas and automatically generate and execute SQL statements, the problem of time-consuming writing SQL statements is solved, and data analysis efficiency is improved.
Patent Information
- Application Number
- CN202411684886.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-22
AI Technical Summary
The prior art has a lot of work to write SQL statements when processing large amounts of data, resulting in low data analysis efficiency.
Provides a data analysis method, by adding controls to the canvas to configure data relationships, automatically generates SQL statements for data analysis, and executes the statement to obtain analysis results.
No need for users to manually write SQL statements, which significantly improves data analysis efficiency and simplifies the data analysis process.
Smart Images

Figure CN119179710B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis, and in particular to a data analysis method, device, storage medium and electronic device. Background Art
[0002] Business data analysis usually requires writing SQL (Structured Query Language, SQL) statements and executing them to obtain data analysis results. However, when there is a lot of data to be analyzed, the workload of writing SQL statements is large and the data analysis efficiency is low. Therefore, how to automatically generate SQL statements for data analysis and improve data analysis efficiency is an urgent problem to be solved.
[0003] Based on this, this specification provides a data analysis method. Summary of the invention
[0004] This specification provides a data analysis method, device, storage medium and electronic device to partially solve the above-mentioned problems existing in the prior art.
[0005] This manual adopts the following technical solutions:
[0006] This specification provides a data analysis method, the method comprising:
[0007] In response to a user's operation of requesting data analysis, presenting a canvas to the user, the canvas including a plurality of controls for configuring data relationships;
[0008] In response to the data adding operation of the user, acquiring the data to be analyzed added by the user in the canvas;
[0009] In response to the user's operation on the control, determining configuration information of the control, the configuration information including an operation condition and an operation object;
[0010] In response to a data analysis trigger operation by a user, determining an SQL statement of the data to be analyzed based on the control and configuration information of the control;
[0011] Execute the SQL statement to obtain the data analysis result of the data to be analyzed and display it.
[0012] Optionally, determining the configuration information of the control specifically includes:
[0013] In the canvas, determining a data analysis result generation node;
[0014] According to the preset relationship attributes, the front control of the data analysis result generation node is determined, and according to the attribute identifier of the front control and the preset corresponding relationship, the configuration information of the front control is determined until the configuration information of all controls operated by the user is obtained.
[0015] Optionally, in the canvas, determining a data analysis result generation node specifically includes:
[0016] Obtaining attribute identifiers of several nodes in the canvas;
[0017] According to the attribute identifier, determining whether there is a data analysis result generation node;
[0018] If not, the node whose attribute is identified as corresponding to the preset identification is determined as the data analysis result generation node.
[0019] Optionally, the control includes at least one of a connection control, a filter control, a grouping control, a function control, and a sorting control.
[0020] Optionally, when the control is a connection control, before determining the SQL statement of the data to be analyzed based on the control and the configuration information of the control, the method further includes:
[0021] Determine whether several operation objects of the connection control are repeated and / or determine whether the operation object of the connection control is missing;
[0022] If yes, a connection error message is returned and displayed.
[0023] Optionally, determining the SQL statement of the data to be analyzed based on the control and the configuration information of the control specifically includes:
[0024] For the query SQL statement, the operation object in the configuration information of the filter control and the operation object in the configuration information of the function control are spliced according to a preset first splicing format to obtain a query SQL statement;
[0025] For the screening SQL statement, determine the screening SQL statement according to the operation conditions in the configuration information;
[0026] For the source SQL statement, determine whether the number of the operation objects of the connection control and the number of the connection controls meet a preset quantity relationship, and if so, connect the operation objects of the connection controls according to a preset second splicing format to obtain a connection SQL statement;
[0027] For the grouping SQL statement, obtain the operation object of the grouping control, and splice the operation objects of the grouping control according to the preset third splicing format to obtain the grouping SQL statement;
[0028] For the sorting SQL statement, for each sorting control, determine whether the sorting control has an operation object according to the configuration information of the sorting control. If so, sort the operation object of the sorting control according to the operation conditions of the sorting control to obtain the sorting SQL statement.
[0029] Optionally, before connecting the operation object of the connection control, the method further includes:
[0030] For each connection control, determining whether the preset first operation object of the connection control is the preset second operation object of the connection control preceding the connection control;
[0031] If not, for each connection control, the operation object of the connection control is replaced with the operation object of each remaining connection control to obtain a plurality of replacement results, and a target replacement result is determined among the plurality of replacement results.
[0032] Optionally, the method further comprises:
[0033] Generate duplicate data based on business data and use it as data to be analyzed;
[0034] When the preset time arrives, judging whether the business data is updated according to the data to be analyzed;
[0035] If so, the data to be analyzed is updated according to the updated business data.
[0036] This specification provides a data analysis device, the device comprising:
[0037] A canvas display module, used for displaying a canvas to a user in response to a user's operation of requesting data analysis, wherein the canvas includes a plurality of controls for configuring data relationships;
[0038] A data acquisition module, configured to acquire the data to be analyzed added by the user in the canvas in response to the user's data adding operation;
[0039] A configuration information determination module, configured to determine configuration information of the control in response to the user's operation on the control, wherein the configuration information includes an operation condition and an operation object;
[0040] An SQL statement determination module, for determining, in response to a data analysis trigger operation of a user, an SQL statement of the data to be analyzed based on the control and configuration information of the control;
[0041] The statement execution module is used to execute the SQL statement, obtain the data analysis results of the data to be analyzed, and display them.
[0042] Optionally, the configuration information determination module is specifically used to determine, in the canvas, a data analysis result generation node; determine a front control of the data analysis result generation node according to preset relationship attributes; and determine the configuration information of the front control according to the attribute identifier of the front control and a preset corresponding relationship, until the configuration information of all controls operated by the user is obtained.
[0043] Optionally, the configuration information determination module is specifically used to obtain attribute identifiers of several nodes in the canvas; based on the attribute identifiers, determine whether there is a data analysis result generation node; if not, determine the node whose attribute identifier is a preset identifier as the data analysis result generation node.
[0044] Optionally, the control includes at least one of a connection control, a filter control, a grouping control, a function control, and a sorting control.
[0045] Optionally, the device further comprises:
[0046] The first verification module is used to determine whether several operation objects of the connection control are repeated and / or whether the operation object of the connection control is missing before determining the SQL statement of the data to be analyzed based on the control and its configuration information when the control is a connection control; if so, return a connection error prompt message and display it.
[0047] Optionally, the SQL statement determination module is specifically used to, for a query SQL statement, splice the operation objects in the configuration information of the filter control and the operation objects in the configuration information of the function control according to a preset first splicing format to obtain a query SQL statement; for a filter SQL statement, determine the filter SQL statement according to the operation conditions in the configuration information; for a source SQL statement, determine whether the number of operation objects of the connection control and the number of connection controls meet a preset quantity relationship; if so, connect the operation objects of the connection control according to a preset second splicing format to obtain a connection SQL statement; for a grouping SQL statement, obtain the operation object of the grouping control, splice the operation objects of the grouping control according to a preset third splicing format to obtain a grouping SQL statement; for a sorting SQL statement, for each sorting control, determine whether the sorting control has an operation object according to the configuration information of the sorting control; if so, sort the operation objects of the sorting control according to the operation conditions of the sorting control to obtain a sorting SQL statement.
[0048] Optionally, the device further comprises:
[0049] The second verification module is used to determine, for each connection control, whether the preset first operation object of the connection control is the preset second operation object of the connection control before connecting the operation objects of the connection control; if not, for each connection control, permuting the position of the operation object of the connection control with the operation objects of each remaining connection control to obtain a plurality of permutation results, and determining a target permutation result among the plurality of permutation results.
[0050] Optionally, the device further comprises:
[0051] The data synchronization module is used to generate duplicate data based on business data and use it as data to be analyzed; when a preset time arrives, it is determined whether the business data is updated based on the data to be analyzed; if so, the data to be analyzed is updated based on the updated business data.
[0052] This specification provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned data analysis method is implemented.
[0053] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned data analysis method when executing the program.
[0054] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:
[0055] It can be seen from the data analysis method provided in this specification that in response to the user's data addition operation, the data to be analyzed can be obtained. In response to the user's operation on the control, configuration information can be determined, and the configuration information can characterize the data relationship between the data to be analyzed. Then, according to the configuration information of the control and the control, an SQL statement is generated, and the SQL statement is executed to obtain the data analysis result. There is no need for the user to write an SQL statement for data analysis, which improves the efficiency of data analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The drawings described herein are used to provide a further understanding of this specification and constitute a part of this specification. The illustrative embodiments and descriptions of this specification are used to explain this specification and do not constitute an improper limitation on this specification. In the drawings:
[0057] Figure 1 A schematic diagram of a data analysis method provided in this specification;
[0058] Figure 2 A schematic diagram of a canvas provided for this specification;
[0059] Figure 3A schematic diagram of the parameter configuration area provided in this manual;
[0060] Figure 4 A schematic diagram showing the data analysis results provided in the manual;
[0061] Figure 5 A schematic diagram of a dictionary table structure provided in this manual;
[0062] Figure 6 A schematic diagram of a field dictionary table structure provided for this specification;
[0063] Figure 7 A schematic diagram of a data analysis device provided in this specification;
[0064] Figure 8 A method corresponding to the Figure 1 Schematic diagram of the structure of an electronic device. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.
[0066] The technical solutions provided by the embodiments of this specification are described in detail below in conjunction with the accompanying drawings.
[0067] When analyzing data, it is usually necessary to write corresponding SQL statements and execute them to obtain data analysis results. Writing SQL statements requires certain writing skills. When the SQL statements to be written are relatively complex and the amount of data to be analyzed is too large, it requires the writers to have high writing skills and takes more time, which greatly reduces the efficiency of data analysis. Therefore, this specification provides a data analysis method. The execution subject of this specification can be an electronic device that can execute the solution of this specification, such as a server, a mobile phone, a personal computer (PC), a distributed cluster, a data engine, etc., and can also be other computing devices with computing capabilities. This specification does not limit this. For the sake of convenience, this specification is described with the server as the execution subject. The terminal can communicate and exchange data with the server, and the user uses the terminal to obtain various types of information sent by the server.
[0068] Figure 1 A flow chart of a data analysis method provided in this specification includes the following steps:
[0069] S100: In response to a user's operation of requesting data analysis, a canvas is displayed to the user, where the canvas includes a plurality of controls for configuring data relationships.
[0070] In one or more embodiments of this specification, a user may request data analysis according to data analysis requirements. Specifically, a user may click on a canvas icon displayed in a terminal or in other forms, so that the server receives a request from the user to analyze the data, which is not limited in this specification. Afterwards, the server provides a canvas and displays the canvas through the terminal used by the user, so that the user can perform operations such as adding data and configuring data relationships in the canvas. It can also be understood that the server displays the canvas to the user in response to the user's operation of requesting data analysis.
[0071] Figure 2 A canvas diagram provided for this specification, such as Figure 2 shown.
[0072] The canvas may include a drawing area, a control selection area, a data storage area to be analyzed, etc. The data storage area to be analyzed is used to store the data to be analyzed, and the data can be imported into the data storage area to be analyzed by the existing data import method. This specification does not limit the type of data to be analyzed. For the convenience of subsequent description, the data in the form of a table is used for description. For example, the data storage area to be analyzed can store the xx table. The user can select several controls in the control selection area to configure data relationships. The drawing area can be used to display the data to be analyzed added by the user from the data storage area to be analyzed, and can also be used to display the data relationship of the data to be analyzed configured by the user using the control. Among them, the control includes at least one of a connection control, a filter control, a grouping control, a function control, a sorting control, an end control, and a table control.
[0073] Of course, the canvas may also include a query area, where users can enter keywords to query the data in the data storage area to be analyzed, so that the queried data can be added to the drawing area later. Figure 2 Only an example of a canvas is provided for this specification. This specification does not limit the specific arrangement of the areas in the canvas, nor does it limit other areas related to data analysis that the canvas may include, such as the canvas may also include an area for displaying data analysis results.
[0074] S102: In response to the data adding operation of the user, obtaining the data to be analyzed added by the user in the canvas.
[0075] Specifically, the user can add the data to be analyzed to the drawing area of the canvas by dragging a data table from the data storage area to be analyzed, clicking a data table displayed in the data storage area to be analyzed, etc. The server responds to the user's data adding operation and obtains the data to be analyzed added by the user in the canvas, that is, obtains the data to be analyzed added by the user in the drawing area of the canvas.
[0076] S104: In response to the user's operation on the control, determining configuration information of the control, the configuration information including an operation condition and an operation object.
[0077] In one or more embodiments of the present specification, the canvas may further include a parameter configuration area. When a user operates a control, the parameter configuration area may be displayed to allow the user to configure data relationships according to personal needs, thereby facilitating subsequent server determination of control configuration information.
[0078] In addition, there are differences in the configuration information of different controls. Specifically, the connection control can be used to connect the data to be analyzed. If the data to be analyzed is in tabular form, the operation object included in the configuration information of the connection control is the left table name and the right table name, and the operation condition includes the left table field name, the right table field name and the operator. Among them, the left table can also be understood as the first operation object, and the right table is the second operation object. The left table and the right table are relative, but the connection order is immutable. For example, the left table of the first connection control is Table 1, and the right table is Table 2. The left table of the second connection control is Table 2, and the right table is Table 3. After the connection, it should be Table 1-Table 2-Table 3. If the left table of the second connection control is Table 3 and the right table is Table 2, it should be Table 1-Table 2, Table 3-Table 2 after the connection, and two tables are obtained, while Table 1-Table 2-Table 3 is one table. Generally speaking, the order of Table 1-Table 2 and Table 3-Table 2 is not feasible, which will be described in detail later and will not be repeated here.
[0079] Operators include greater than, equal to, less than, greater than or equal to, less than or equal to, not equal to, etc. The configuration information of the connection control also includes the connection type, which includes inner join (INNER JOIN): only returns the data that meets the connection conditions in the two tables, left join (LEFT JOIN): returns all the data in the left table, even if there is no data in the right table that meets the connection conditions, right join (RIGHT JOIN): returns all the data in the right table, even if there is no data in the left table that meets the connection conditions, etc. This manual will not list them one by one here.
[0080] Figure 3 This is a schematic diagram of the parameter configuration area provided in this manual, such as Figure 3 shown.
[0081] The FROM clause in SQL query can be generated by using the connection control and its configuration information, especially when it involves multiple table connections. Take the connection control as an example. There are Table 1 and Table 2. Using the connection control to configure Table 1 and Table 2 indicates that Table 1 and Table 2 are connected, that is, the relationship between Table 1 and Table 2 is a connection relationship. Among them, the operation objects of the connection control are Table 1 and Table 2, the left table is Table 1, the right table is Table 2, the operator in the operation condition is equal, and the connection type is inner connection. Then, when the field name in Table 1 is equal to the field name in Table 2, the data with the same field name in Table 1 and Table 2 are connected to obtain the data analysis result. If the field name is "Department ID", the data in Table 1 is: Name: Xiaohong, Department ID: 10, and the data in Table 2 is Department ID: 10, Department Name: Life Department, then the data analysis result is Xiaohong, Life Department.
[0082] The filter control can be used to filter the data to be analyzed and obtain the data that meets the operation conditions. If the data to be analyzed is in the form of a table, the filter control includes a column filter control for the SELECT clause table and field and a row filter control for the WHERE clause query condition. The configuration information of the column filter control is used to generate the SELECT clause in the SQL query to determine the query object, which is usually stored in the form of a linked list. The operation object of the configuration information of the column filter control includes the table name, and the operation condition includes the filter field corresponding to the table name, which is used for the splicing of the SELECT statement. For example: to filter the data of the field col1 in table1, the operation object is table1, and the filter fields are col1 and col2; to filter the data of the field col3 in table2, the operation object is table2, and the filter field is col3, the final result is as follows: SELECT table1.col1, table1.col2, t2.col3. The configuration information of the row filter is used to generate the WHERE clause in the SQL query to filter data according to specific conditions. The operation object includes the table name, and the operation condition includes the operation object (table name), operator, filter field, filter data value and the type of filter data value. The type of filter data value includes character type, integer type, etc. The configuration information of the row filter control is usually stored in the form of nested lists. Each inner list represents a set of operation conditions. Multiple inner lists are connected by logical OR, and the operation conditions in the same group are connected by logical AND. For example, the first set of operation conditions: table name: table1, filter field: col1, operator: = (equal to), filter data value: value1, data type: string, table name: table1, filter field: col2, operator: > (greater than), filter data value: 10, data type: integer. The second set of operation conditions: table name: table2, filter field: col3, operator: <> (not equal to), filter data value: value3, data type: string. The generated WHERE clause: where (table1.col1 = 'value1' and t1.col2>10) or (t2.col3<>'value3').
[0083] The grouping control can be used to generate the GROUP BY clause in the SQL query to group the data to be analyzed. The operation object includes the table name, and the operation condition includes the field name. For example, Table 1 includes the names of several branches of a chain supermarket x in various cities, City 1, Branch Supermarket 1, City 2, Branch Supermarket 2, City 1, Branch Supermarket 3, City 2, Branch Supermarket 4. Table 1 can be grouped according to the city name, and the data analysis results obtained are City 1, Branch Supermarket 1, City 1, Branch Supermarket 3, City 2, Branch Supermarket 2, City 2, Branch Supermarket 4.
[0084] The function control can be used to use several preset functions to perform function calculations on the data to be analyzed. The configuration information of the function control includes the operation object as the table name, and the operation conditions include the function name and field name. The preset function can be set as needed, and this manual does not limit the specific function.
[0085] The sort control is used to sort the data to be analyzed. The operation object includes the table name, and the operation conditions include the field name, descending sort, and ascending sort. The end control is used to end the data analysis.
[0086] After the user operates the control, the server may determine the configuration information of the control in response to the user's operation on the control.
[0087] S106: In response to a data analysis trigger operation by a user, determining an SQL statement for the data to be analyzed based on the control and configuration information of the control.
[0088] After the user adds the data to be analyzed according to the needs and configures the data relationship through the control, the data analysis can be triggered. The server responds to the user's data analysis trigger operation and determines the SQL statement of the data to be analyzed based on the control and its configuration information.
[0089] Since the user has added the data to be analyzed in the drawing area and configured the relationship between the data to be analyzed through the control, the server can generate the corresponding SQL statement based on the obtained configuration information. This is because the SQL statement is mainly composed of five parts: query (SELECT), filter (WHERE), source (FROM), grouping (GROUP BY), and sorting (ORDER BY). The specific parameters required for the content included in the aforementioned SQL statement are included in the configuration information. The server can obtain the corresponding SQL statement by determining the position of the operation object and operation condition in the SQL statement in the configuration information, without the need for the user to write the SQL statement.
[0090] In addition, for salesmen without writing ability (a type of user), since salesmen have a better understanding of the actual business, the data to be analyzed selected by the salesmen have a greater impact on the subsequent execution of the business. The salesmen need to inform the writing personnel (a type of user) of the data analysis requirements, and the writing personnel will write SQL statements to analyze the data. In this manual, salesmen can directly perform data analysis through the canvas provided by the server, and no longer need to obtain data analysis results through writing personnel, which simplifies the data analysis process, makes data analysis more convenient and low-threshold, and thus improves data analysis efficiency.
[0091] S108: Execute the SQL statement to obtain the data analysis result of the data to be analyzed, and display it.
[0092] The canvas may also include a display area for displaying data analysis results. After the server obtains the SQL statement, it can execute the SQL statement to obtain the data analysis results of the data to be analyzed, and display the data analysis results in the display area of the canvas, so that the data analysis results are visualized and more intuitive.
[0093] Figure 4 A schematic diagram showing the data analysis results provided in the manual, such as Figure 4 shown.
[0094] The display area can display the column name, data type of each column data, etc.
[0095] based on Figure 1 The data analysis method shown can obtain the data to be analyzed in response to the user's data addition operation. In response to the user's operation on the control, configuration information can be determined, and the configuration information can characterize the data relationship between the data to be analyzed. Then, according to the configuration information of the control and the control, an SQL statement is generated, and the SQL statement is executed to obtain the data analysis result. There is no need for the user to write an SQL statement for data analysis, which improves the efficiency of data analysis.
[0096] With respect to step S104, when determining the configuration information of the control, since the user may operate more than one control, in order not to miss each control, the controls operated by the user in the canvas may be traversed to obtain the configuration information of each control.
[0097] Specifically, during the traversal process, the server may first determine the data analysis result generation node in the canvas, from which the data analysis result is generated. Then, according to the preset relationship attribute, the front control of the data analysis result generation node is determined, and according to the attribute identifier of the front control and the preset corresponding relationship, the configuration information of the front control is determined until the configuration information of all controls operated by the user is obtained.
[0098] The preset relationship attributes include the relationship source and the relationship result. For example, for existing Table 1 and Table 2, the result 1 is obtained by connecting the control, and the result 2 is obtained by filtering the result 1 using the filter control. Then, the result 2 is analyzed by the grouping control to obtain the result 3, and the data analysis ends. For the connection control, since the connection control has no preceding control, there is no relationship source, and the relationship result is the filter control. For the filter control, the relationship source is the connection control, and the relationship result is the grouping control. For the grouping control, the relationship source is the filter control, and the relationship result is the result control. Therefore, the preceding control of each node can be determined according to the preset relationship attributes to avoid missing controls.
[0099] Furthermore, the control also has an attribute identifier, which can be used to distinguish the control type and determine the configuration information of the control. This is because the attribute identifiers of different types of controls are different, and there is a correspondence between the attribute identifier and the configuration information. During the configuration process, the user will also improve the configuration information. For example, the attribute identifier of the grouping control is Group, and the configuration information includes the table name and field name. After the user completes the configuration, the table name in the configuration information is Table 3 and the field name is Age. Therefore, the server can determine the configuration information of the front control through the attribute identifier of the front control and the preset correspondence. Then, determine the front control of the front control, and determine the configuration information, until the configuration information of all controls operated by the user is obtained.
[0100] When obtaining configuration information, if the control attribute is marked as "Order", the control is a sorting control. The server can obtain the configuration information of the sorting control through JSON escape, which is List <orderitemconfig>orderItemConfigs.
[0101] If the control attribute is marked as "Group", the control is a group control. The configuration information of the classification control is obtained through JSON escape. GroupConfig groupConfig includes: Group parameter list List <groupconditionconfig>conditions, each item in the linked list contains the table name and field name.
[0102] If the control attribute is marked as "Filter", the control is a filter control. The configuration of the filter control is obtained through JSON escape. The configuration information of the column filter control includes: Column filter parameter list List <filtercolconfig>filterColConfigs. For example, if the SELECT clause is select table1.col1, table1.col2,table2.col3, the column filter configuration is as follows: List <filtercolconfig>colConfigs =Arrays.asList new FilterColConfig("table1", new String[]{"col1", "col2"}),new FilterColConfig("table2", new String[]{"col3"})).
[0103] The configuration information of the row filter control includes: row filter parameter list List <List <filterrowconfig>>filterRowConfigsList. For example, if the WHERE clause is WHERE (t1.col1 = 'value1' and t1.col2>10) or (t2.col3<>'value3'), the row filter configuration is as follows: List <List <filterrowconfig>>rowConfigsList = Arrays.asList(Arrays.asList(new FilterRowConfig("
[0104] table1","col1", OperatorType.EQUAL, "value1",TableField.DataType.STRING), new FilterRow
[0105] Config("table1", "col2", OperatorType.GREATER_THAN, "10",TableField.DataType.INTEGER
[0106] )), the first inner list: contains two operation conditions of the second configuration information FilterRowConfig, both of which are for the table named table1. The first operation condition means applying the equal operator to the col1 column, the value is value1, and the data type is a string. The second operation condition means applying the greater than operator to the col2 column, the value is 10, and the data type is an integer.
[0107] Arrays.asList(new FilterRowConfig("table2", "col3", OperatorType.NOT_EQUAL, "value3", TableField.DataType.STRING))). The second internal list: contains an operation condition of the second configuration information FilterRowConfig, which is for the table named table2. This operation condition means applying the not equal operator to the col3 column, the value is value3, and the data type is string.
[0108] If the control attribute is marked as "Connect", the control is a connection control. The configuration information of the connection control is obtained through JSON escape, which is a List <connectconfig>connectConfigs. For example, the FROM clause is: from table1 t1inner join table2 t2 on t1.col1 = t2.col2 left join table3 t3 on t2.col3 = t3.col4, then the connection filter configuration is as follows: List <connectconfig>connectConfigs = Arrays.asList(
[0109] new ConnectConfig("table1", "table2", "inner join", Arrays.asList(
[0110] new ConnectOnConfig("col1", "col2", OperatorType.EQUAL)
[0111] )),
[0112] new ConnectConfig("table2", "table3", "left join", Arrays.asList(
[0113] new ConnectOnConfig("col3", "col4", OperatorType.EQUAL)
[0114] ))). The configuration information of the first ConnectConfig connection control: left table: table1, right table: table2, connection type: inner join, connection condition: left table column name: col1, right table column name: col2, operator: OperatorType.EQUAL equals. The configuration information of the second ConnectConfig connection control: left table: table2, right table: table3, connection type: left join, connection condition: left table column name: col3, right table column name: col4, operator: OperatorType.EQUAL equals. First, an inner join is performed between table1 and table2, and then a left join is performed between table2 and table3.
[0115] If the control attribute is marked as "TableConfig", the control is a table control. The configuration information of the table control to be analyzed is obtained through JSON escape, which is a List <tableconfig>tableConfigs. For example, List <tableconfig>tableConfigs = Arrays.asList(
[0116] new TableConfig("table1", "t1"),
[0117] new TableConfig("table2", "t2")
[0118] ); that is, you need to query table1, alias t1 and table2, alias t2. The alias of table1 is t1 and the alias of table2 is t2.
[0119] The aforementioned attribute identification is an embodiment provided in this specification, which can be set specifically as needed, and this specification does not limit this.
[0120] It should be noted that the controls used by the user can also constitute nodes in the drawing area, and the data to be analyzed does not constitute a node. The above result 2 can be understood as an intermediate result, and the direction of the edge drawn by the user in the drawing area is the direction of data flow. Figure 2 As shown, there are two nodes in the drawing area, namely, a node consisting of a connection control and a node consisting of an end control. The edges of Table 1 and Table 2 point to the connection control, indicating that Table 1 and Table 2 flow to the connection control. Then, when the server determines the data analysis result generation node, it can first obtain the attribute identifiers of several nodes in the canvas, and determine whether there is a data analysis result generation node based on the attribute identifier. If there is a data analysis result generation node, it indicates that the data analysis result viewed by the user is an intermediate result. Specifically, it can be determined whether there is a node with an attribute identifier of executeId. If so, the node with the attribute identifier of executeId is determined as a data analysis result generation node. If not, it means that the user may have accidentally deleted or forgotten to specify the data analysis result generation node. In order to enable the data analysis result to be generated normally, the node with the attribute identifier corresponding to the preset identifier can be determined as a data analysis result generation node. The preset identifier can be the attribute identifier of the end control, such as end. executeId is only an example of an attribute identifier for specifying a data analysis result generation node, and this manual does not limit this.
[0121] Before executing step S106, when the control is a connection control, the server may also verify the configuration information of the connection control to determine whether the operation objects of the connection control are repeated and / or whether the operation object of the connection control is missing. If so, a connection error prompt message is returned and displayed. The server may only determine whether the operation objects of the connection control are repeated, or only determine whether the operation object of the connection control is missing, or both determine whether the operation objects of the connection control are repeated and whether the operation object of the connection control is missing, and this specification does not limit this.
[0122] The duplication between operation objects indicates that the table names are repeated. It is possible that the left table and the right table are the same table, and no connection operation is required. The configuration information of the connection control includes at least two operation objects. If there is only one operation object, it means that the operation object is missing and the connection cannot be made. Therefore, the missing operation object will also affect the data analysis results. The server can also determine whether the operation conditions exist. If not, the connection cannot be made. It can also return a connection error prompt message and display it. Of course, when executing step S106, it can also be determined whether there are duplications between several operation objects of the connection control and / or whether the operation object of the connection control is missing.
[0123] For step S106, when generating SQL statements, the server can generate SQL statements for each control using the configuration information of each control. Since the SQL statement mainly consists of five parts: query (SELECT), filter (WHERE), source (FROM), grouping (GROUP BY), and sorting (ORDER BY), the server can determine the specific parameters of these five parts in the configuration information of each control.
[0124] For the query SQL statement, i.e., the SELECT clause, the operation object in the configuration information of the column filter control and the operation object in the configuration information of the function control are spliced according to the preset first splicing format to obtain the query SQL statement. The preset first splicing format is select operation object 1, operation object 2, ...
[0125] For example, if the operation object in the filter control is the col1 field of Table 4, the operation object in the function control is the col1 field of Table 4, and the operation function is MAX, then the query SQL statement is select MAX (Table 4.col1), Table 5. Of course, if the operation object is a column of the table, then the query SQL statement is select Table 4, col1 field, Table 5, col2 field, to query the column data corresponding to the col1 field of Table 4 and the column data corresponding to the col2 field of Table 5.
[0126] In addition, you can also determine whether there are identical fields in the configuration information of the column filter control and the configuration information of the function control. If so, rename the identical fields to avoid duplication of identical fields in different tables in the SELECT clause. Specifically, you can define a Map<String, Integer> fieldCountMap = new HashMap<>(), stores fields. If the currently stored fields already exist in fieldCountMap, it means that fields in different tables have duplicate names. Therefore, the same fields in different tables can be renamed. If the column filter control configuration information is empty, it means that no fields are specified for query, and all fields in the table can be queried, that is, SELECT *.
[0127] For the screening SQL statement, that is, the WHERE clause, the screening SQL statement is determined according to the operation conditions in the configuration information. Specifically, the operation conditions include screening conditions, operators, etc., which can clarify the user's screening requirements, such as determining data greater than 1 in Table 2.
[0128] According to the configuration information rowConfigsList of the row filter control, traverse rowConfigsList, determine the operation conditions, feel the operation conditions, and determine the filter SQL statement. Among them, a first configuration information rowConfig can still constitute a list list to obtain rowConfigsList. Each first configuration information rowConfigsList may include several second configuration information filterRowConfigs. The second configuration information filterRowConfig contains the table name and the column name of the table. Multiple filterRowConfigs are associated with "and", and multiple configuration rowConfigs are associated with "or". The above-mentioned examples already exist, so they will not be repeated here.
[0129] For the source SQL statement, i.e., the FROM clause, it is determined whether the number of the operation objects of the connection control and the number of the connection controls meet the preset quantity relationship. If so, the operation objects of the connection control are connected according to the preset second splicing format to obtain the connection SQL statement. The preset second splicing format is that the left table name is in front and the right table name is in the back.
[0130] Specifically, determine whether the number of operation objects of the connection control is empty. If so, it means that there is no multi-table connection relationship and it is a single-table query. If not, it is a joint table query and needs to be verified. The server can first determine whether the number of tables is the number of connection controls + 1. If not, the table connection relationship is incorrect and a replacement operation is performed. The details have been described later and will not be repeated here. If correct, recursively check whether the order of the connection controls is feasible to form a complete connection table relationship. After obtaining a complete connection table relationship, when the operation conditions of the connection control are met, the operation objects are connected through conditional statements such as left connection, right connection, inner connection, etc.
[0131] Is the order of the connection controls feasible? If it is feasible, for example, if there are the following two connection control configuration information: Configuration information 1: left table name A, right table name B, connection type: inner connection. Configuration information 2: left table name B, right table name C, connection type: inner connection. First check the initial state: [ConnectConfig(A, B), ConnectConfig(B, C)], check the order: traverse the configuration information of the connection controls, and check whether the left table name and right table name of each connection control are in the table name set that has appeared. The configuration information of the first connection control ConnectConfig(A, B): the left table name A and the right table name B are both new, and are added to the set. The configuration information of the second connection control ConnectConfig(B, C): the left table name B is already in the set, and the right table name C is new, and is added to the set. Result: The left table name and right table name of all connection controls are in the table name set that has appeared, so this order is feasible, and the query SQL statement can be determined.
[0132] Infeasible situation: If the following two connection control configuration information exist: Configuration information 1: left table name B, right table name C, connection type: inner connection. Configuration information 2: left table name A, right table name B, connection type: inner connection. Initial state: Connection control configuration information: [ConnectConfig(B, C), ConnectConfig(A, B)], verification order: traverse the configuration information of each connection control, and check whether the left table name and right table name of each connection control are in the table name set that has appeared. The first connection control ConnectConfig(B, C): The left table name B and the right table name C are both new, and are added to the set. The second connection control ConnectConfig(A, B): The left table name A is new, the right table name B is already in the set, but the left table name A is not in the set, so this order is not feasible, and an error is returned.
[0133] For the grouping SQL statement, that is, the GROUP BY clause, the operation object of the grouping control is obtained, and the operation objects of the grouping control are spliced according to the preset third splicing format to obtain the grouping SQL statement. The preset third splicing format is the table name.column name format, and multiple "table name.column name" are separated by ",".
[0134] Specifically, the server can traverse the configuration information linked list groupConfig according to the configuration information linked list of the grouping control, obtain the table name and column name in each linked list, and splice them into the form of "table name.column name", and multiple "table name.column name" are separated by ",".
[0135] For the sorting SQL statement, i.e., the order by clause, for each sorting control, it is determined whether the sorting control has an operation object according to the configuration information of the sorting control. If so, the operation object of the sorting control is sorted according to the operation conditions of the sorting control to obtain the sorting SQL statement.
[0136] Specifically, the server determines whether the number of operation objects is 0 according to the operation object of the sort control. If it is 0, there is no sort object, and an error message can be returned to prompt the user that there is an error in the sort. If it is not empty, a sort combination is performed. Specifically, the table name, column name, and whether it is in reverse or forward order can be obtained by traversing the configuration information of the sort control. Different operation conditions are separated by "".
[0137] Before connecting the operation object of the connection control, the configuration information of the connection control needs to be recursively checked again, in order to recursively check whether the order of the connection controls is feasible to form a complete connection table relationship.
[0138] The steps of the server verifying the connectivity of the configuration information of the connection control and performing sequence optimization are as follows: creating an operation object repository, determining the execution order of the connection control, traversing the configuration information of each connection control according to the execution order, and storing the operation object of the connection control in the operation object repository for each connection control. In the operation object repository, it is determined whether the preset first operation object of the connection control is the preset second operation object of the preceding connection control of the connection control. That is, whether the left table of the connection control is the right table of the preceding connection control of the connection control. If not, it means that the table connection is incorrect, then for each connection control, the operation object of the connection control is permuted with the operation object of each remaining connection control to obtain a number of permutation results, and a target permutation result is determined among the several permutation results, and the target permutation result satisfies that the preset first operation object of the connection control is the preset second operation object of the preceding connection control of the connection control.
[0139] The server can also perform parameter verification first. If the number of operation objects of the connection control is 0 or there is only one operation object, an error message is directly returned and displayed to the user without further verification and optimization. Otherwise, the server can verify the connectivity and order optimization of the configuration information of the connection control. It can be understood that when performing connectivity verification, the server can use the Set collection to store all table names, traverse the configuration information of the connection control, and add the left table name and right table name of each connection control to the established table name storage set tableNameSet. Determine whether the number of table names in tableNameSet is equal to the number of connection controls plus one. If not, it means that the number of tables does not match the number of connection controls, the table connection relationship is incorrect, and an exception message is returned to prompt the user that there is a problem with the current connection.
[0140] When performing sequential optimization, the server can first determine whether the number of connectConfigs linked lists is greater than a preset number threshold. For example, if the preset number threshold is 2, it determines whether the number of connected controls is greater than the preset number threshold. If so, the optimization function boolean optimize(List <connectconfig>connectConfigs, int index), recursively adjust the order of connection controls to ensure that they can form a complete connection table relationship. Among them, index is the position of the current traversal.
[0141] In order to determine the correct connection and adjust the order of the connection controls, for each connection control, the operation object of the connection control is replaced with the position of the operation object of each connection control. The purpose of the replacement is to try all feasible linked list sequences of the connection controls, so as to find a valid sequence that can satisfy the link list relationship. Specifically, the replacement operation systematically explores all possible permutations and combinations through recursion and backtracking to ensure that a feasible sequence can be found or that there is no feasible sequence.
[0142] Specifically, when index is the last element in the connection control connectConfigs, it means that the last connection control has been reached. In this case, it is necessary to check whether the linked list order of the current connection control is feasible. The server can initialize a Set collection to store the table names that have appeared, traverse connectConfigs, and start from the second connection control (i.e., i = 1) to check whether the left table name and right table name of each connection control are in the set of table names that have appeared. If the left table name and the right table name of the connection control are not in the set, it means that the operation object of the left connection or right connection of the connection control is missing. When index does not reach the last element of connectConfigs, traverse connectConfigs, start from the current index index, exchange the connection control positions of the current index index and the subsequent index i, recursively call the optimize(connectConfigs, index+ 1) method, and determine whether the linked list order of the connection control starting from the index+ 1 position is correct. If correct, determine the query SQL statement. Otherwise, backtrack to restore the state before the exchange and continue to try other possible orders.
[0143] For example, if there are three connection controls, the corresponding configuration information ConnectConfig list: Configuration information 1 ConnectConfig 1: leftTableName = "A", rightTableName = "B". Configuration information 2 ConnectConfig 2: leftTableName = "B", rightTableName = "C". Configuration information 3 ConnectConfig 3: leftTableName = "C", rightTableName = "D". Call the optimization function boolean optimize(List <connectconfig>connectConfigs, int index), adjust the order of the aforementioned connection controls to ensure that they can form a valid connection table relationship.
[0144] Specific process: Initial call: optimize(connectConfigs, 0), index = 0, did not reach the last element, entered the loop. Try to swap the connection controls at positions 0 and 0, call optimize(connectConfigs,1). First recursion: optimize(connectConfigs, 1), index = 1, did not reach the last element, entered the loop. Try to swap the connection controls at positions 1 and 1, call optimize(connectConfigs, 2). Second recursion: optimize(connectConfigs, 2), index = 2, reached the last element, and checked. Initialize Set, store the table names that have appeared: {"A", "B", "C", "D"} Traverse the above three connection controls: The first connection control: leftTableName = "A", rightTableName = "B", are all in the set. The second connection control: leftTableName = "B", rightTableName = "C", are all in the set. The third connection control: leftTableName = "C", rightTableName = "D", are all in the set. The order of connecting controls is correct, which determines that the intermediate recursive result is correct: optimize([1, 2, 3], 1), and the final result: optimize([1, 2, 3],0).
[0145] It can be understood that when determining the SQL statement, in the query part, first determine whether there is configuration information for the column filter control. If so, query the specified field. If not, query all fields of the table by default. If there is configuration information for the function control, determine whether the field of the table needs to be processed by the function. The fields are separated by "", such as MAX(ID). In the filter part, operators and constants are spliced from the configuration information of the row filter control. Fields are separated by "OR", such as ID>5. Note that if the constant value is a character type, double quotes need to be added. In the source part, splicing is done through the configuration information of the connection control, such as "FROM table1 JOIN table2 ON table1 ON table1.a = table2.a AND table1.b =table2.b", and the connection conditions after ON are separated by "AND". In the sorting part, the table name and field name are assembled through the configuration information of the grouping control, such as "GROUP BY table1.a", and the connection conditions after "GROUP BY" are separated by ",". The sorting part is assembled through the configuration information of the sorting control, such as "ORDER BY table1.a DESC", "the conditions after ORDER BY are separated by ",". After the above assembling is completed, it is reassembled into a complete SQL statement, and the generated SQL statement is executed according to the preset relationship attributes. In addition, the execution control of the canvas will also be saved. When the canvas is closed and opened again, the front end re-parses and can still be reproduced and continued.
[0146] It can also be understood that for step S108, the server can execute the generated SQL statement according to the preset relationship attributes. This is because each control will generate a corresponding SQL statement, and there is an execution order between the controls. In order to ensure that the final data analysis results are correct, they need to be executed in order, and this order can be determined by the relationship attributes between the controls. For example, there are three controls, namely control 1, control 2 and control 3. Control 3 is the end control. Control 1 has no relationship source, and the relationship result is control 2. The relationship source of control 2 is control 1, and the relationship result is the end control. Therefore, the execution order is to execute the SQL statement of control 1 first, then execute the SQL statement of control 2, and the end control has no SQL statement.
[0147] It should be noted that in order to reduce the risk of leakage of original business data, business data and data to be analyzed are generally stored in different databases. Therefore, the server needs to synchronize the data to be analyzed periodically. Specifically, the server can generate duplicate data based on the business data and use it as the data to be analyzed. When the preset time arrives, it is determined whether the business data is updated based on the data to be analyzed. If so, the data to be analyzed is updated based on the updated business data so that the data to be analyzed can be analyzed again later.
[0148] The data to be analyzed may include a dictionary table table_dict and a field dictionary table table_field_dict. Figure 5 A schematic diagram of a dictionary table structure provided in this manual, such as Figure 5 shown.
[0149] The dictionary table includes fields, types, lengths, whether they are empty, comments, etc. Fields include id, table_name, filed_name, description, data_type, and types include bigint, varchar, text, and tinyint. Comments include field name, field description, data type, and corresponding numbers, 1- string, 2- numeric value, 3- date and time, 4- date.
[0150] Figure 6 A schematic diagram of a field dictionary table structure provided for this specification is as follows: Figure 6 shown.
[0151] Including fields, types, length, whether empty, comments, decimal point, key value, etc. Fields include id, table_name, table_cn_name, description, system_belong, and types include int, varchar, etc. Comments include table name, Chinese table name, Chinese description, name of the system to which it belongs, etc.
[0152] In addition, due to the environmental isolation problem of table_dict, it cannot be read directly. Therefore, it is necessary to create an intermediate field table table_dict_mid. The table structure is consistent with table_dict and is initially empty. The source table table_dict is imported into table_dict_mid for field-related add, delete, modify and query operations.
[0153] Business data can be updated at dawn every day, so set a daily schedule to synchronize the data tables and update the table dictionary table_dict and field dictionary table_field_dict. Specifically, update the table_dict_mid table every day, compare the table_dict_mid business data with the table_dict table, and if the business data does not exist in table_dict_mid, update the business data to the table_dict_mid table. Similarly, compare the table in database db1 with the table_field_dict dictionary table and update the table_field_dict_mid table.
[0154] In addition, users can also create personal control templates in the canvas and manage personal canvases based on templates. The canvas capabilities include but are not limited to creating personal canvas templates, copying templates, saving templates, dragging controls, clicking to execute, viewing execution results, and exporting execution results. The configured canvas template is only for the creator to use, and others have no right to view or modify it. The modified canvas is saved in real time. After closing and reopening, the last content can be reproduced and the above steps S102~108 can be continued. For example, the user can select the corresponding business table, drag controls, splice controls, and configure control information.
[0155] The server can also write business data into a data engine that can be used for Chinese localization, such as ElasticSearch, through data synchronization, and localize the business data through Chinese-English conversion. Users can perform associated searches for Chinese characters on business data.
[0156] The above is one or more implementation methods of this specification, based on Figure 1 The data analysis method shown in the flowchart is provided in this specification. The corresponding data analysis device is also provided. Figure 7 shown.
[0157] Figure 7 A schematic diagram of a data analysis device provided in this specification includes:
[0158] A canvas display module 700, for displaying a canvas to a user in response to a user's operation requesting data analysis, wherein the canvas includes a plurality of controls for configuring data relationships;
[0159] The data acquisition module 702 is used to acquire the data to be analyzed added by the user in the canvas in response to the user's data adding operation;
[0160] A configuration information determination module 704, configured to determine configuration information of the control in response to the user's operation on the control, the configuration information including an operation condition and an operation object;
[0161] An SQL statement determination module 706, for determining, in response to a user's data analysis trigger operation, an SQL statement of the data to be analyzed based on the control and the configuration information of the control;
[0162] The statement execution module 708 is used to execute the SQL statement, obtain the data analysis result of the data to be analyzed, and display it.
[0163] Optionally, the configuration information determination module 704 is specifically used to determine, in the canvas, a data analysis result generation node; determine a front control of the data analysis result generation node according to preset relationship attributes; and determine the configuration information of the front control according to the attribute identifier of the front control and a preset corresponding relationship, until the configuration information of all controls operated by the user is obtained.
[0164] Optionally, the configuration information determination module 704 is specifically used to obtain attribute identifiers of several nodes in the canvas; determine whether there is a data analysis result generation node based on the attribute identifier; if not, determine the node corresponding to the attribute identifier as a preset identifier as the data analysis result generation node.
[0165] Optionally, the control includes at least one of a connection control, a filter control, a grouping control, a function control, and a sorting control.
[0166] Optionally, the device further comprises:
[0167] The first verification module is used to determine whether several operation objects of the connection control are repeated and / or whether the operation object of the connection control is missing before determining the SQL statement of the data to be analyzed based on the control and its configuration information when the control is a connection control; if so, return a connection error prompt message and display it.
[0168] Optionally, the SQL statement determination module is specifically used to, for a query SQL statement, splice the operation objects in the configuration information of the filter control and the operation objects in the configuration information of the function control according to a preset first splicing format to obtain a query SQL statement; for a filter SQL statement, determine the filter SQL statement according to the operation conditions in the configuration information; for a source SQL statement, determine whether the number of operation objects of the connection control and the number of connection controls meet a preset quantity relationship; if so, connect the operation objects of the connection control according to a preset second splicing format to obtain a connection SQL statement; for a grouping SQL statement, obtain the operation object of the grouping control, splice the operation objects of the grouping control according to a preset third splicing format to obtain a grouping SQL statement; for a sorting SQL statement, for each sorting control, determine whether the sorting control has an operation object according to the configuration information of the sorting control; if so, sort the operation objects of the sorting control according to the operation conditions of the sorting control to obtain a sorting SQL statement.
[0169] Optionally, the device further comprises:
[0170] The second verification module is used to determine, for each connection control, whether the preset first operation object of the connection control is the preset second operation object of the connection control before connecting the operation objects of the connection control; if not, for each connection control, permuting the position of the operation object of the connection control with the operation objects of each remaining connection control to obtain a plurality of permutation results, and determining a target permutation result among the plurality of permutation results.
[0171] Optionally, the device further comprises:
[0172] The data synchronization module is used to generate duplicate data based on business data and use it as data to be analyzed; when a preset time arrives, it is determined whether the business data is updated based on the data to be analyzed; if so, the data to be analyzed is updated based on the updated business data.
[0173] This specification also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A data analysis method is provided.
[0174] This manual also provides Figure 8 The one shown corresponds to Figure 1 Schematic diagram of the structure of the electronic equipment. Figure 8 As shown, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 8 The data analysis method described.
[0175] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the executor of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0176] In the 1990s, it was very clear whether the improvement of a technology was a hardware improvement (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or a software improvement (improvement of the method flow). However, with the development of technology, many improvements in the method flow today can be regarded as direct improvements in the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and produce dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0177] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.
[0178] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0179] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0180] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0181] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0182] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0183] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0184] In a typical configuration, a computing device includes one or more processors (first chips), input / output interfaces, network interfaces, and memory.
[0185] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0186] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0187] It should also be noted that the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the existence of other identical elements in the process, method, commodity or device including the element.
[0188] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0189] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, controls, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0190] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0191] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.< / connectconfig> < / connectconfig> < / tableconfig> < / tableconfig> < / connectconfig> < / connectconfig> < / filterrowconfig> < / filterrowconfig> < / filtercolconfig> < / filtercolconfig> < / groupconditionconfig> < / orderitemconfig>
Claims
1. A data analysis method, characterized in that: The method comprises: In response to a user's operation of requesting data analysis, presenting a canvas to the user, the canvas including a plurality of controls for configuring data relationships; In response to the data adding operation of the user, acquiring the data to be analyzed added by the user in the canvas; In response to the user's operation on the control, determining the configuration information of the control, the configuration information including the operation condition and the operation object, wherein in the canvas, determining the data analysis result generation node; determining the front control of the data analysis result generation node according to the preset relationship attribute, and determining the configuration information of the front control according to the attribute identifier of the front control and the preset corresponding relationship, until the configuration information of all controls operated by the user is obtained, the preset relationship attribute including the relationship source and the relationship result; In response to the user's data analysis trigger operation, based on the control and the configuration information of the control, determine the SQL statement of the data to be analyzed, wherein, when the control is a connection control, based on the control and the configuration information of the control, before determining the SQL statement of the data to be analyzed, determine whether several operation objects of the connection control are repeated and / or determine whether the operation object of the connection control is missing, if so, return a connection error prompt message and display it, wherein, when index is the last element in the connection control connectConfigs, initialize a Set collection to store the table names that have appeared, traverse connectConfigs, and check whether the left table name and the right table name of each connection control are in the table name set that have appeared from the connection control indexed by i=1, if the left table name and the right table name of the connection control are not in the set, determine the operation object of the left connection or right connection of the missing connection control, when index does not reach the last element of connectConfigs, traverse connectConfigs, start from the current index index, exchange the connection control positions of the current index index and the subsequent index i, recursively call optimize(connectConfigs, index+ 1) method, and determine from index+ Check whether the linked list sequence of the connection control starting from position 1 is correct. If it is correct, determine the query SQL statement. Otherwise, backtrack to restore the state before the exchange. Execute the SQL statement to obtain the data analysis result of the data to be analyzed and display it.
2. The method according to claim 1, characterized in that In the canvas, determining a data analysis result generation node specifically includes: Obtaining attribute identifiers of several nodes in the canvas; According to the attribute identifier, determining whether there is a data analysis result generation node; If not, the node whose attribute is identified as corresponding to the preset identification is determined as the data analysis result generation node.
3. The method according to claim 1, characterized in that The control includes at least one of a connection control, a filter control, a grouping control, a function control, and a sorting control.
4. The method according to claim 3, characterized in that Determining the SQL statement of the data to be analyzed based on the control and the configuration information of the control specifically includes: For the query SQL statement, the operation object in the configuration information of the filter control and the operation object in the configuration information of the function control are spliced according to a preset first splicing format to obtain a query SQL statement; For the screening SQL statement, determine the screening SQL statement according to the operation conditions in the configuration information; For the source SQL statement, determine whether the number of the operation objects of the connection control and the number of the connection controls meet a preset quantity relationship, and if so, connect the operation objects of the connection controls according to a preset second splicing format to obtain a connection SQL statement; For the grouping SQL statement, obtain the operation object of the grouping control, and splice the operation objects of the grouping control according to the preset third splicing format to obtain the grouping SQL statement; For the sorting SQL statement, for each sorting control, determine whether the sorting control has an operation object according to the configuration information of the sorting control. If so, sort the operation object of the sorting control according to the operation conditions of the sorting control to obtain the sorting SQL statement.
5. The method according to claim 4, characterized in that Before connecting the operation object of the connection control, the method further includes: For each connection control, determining whether the preset first operation object of the connection control is the preset second operation object of the connection control preceding the connection control; If not, for each connection control, the operation object of the connection control is replaced with the operation object of each remaining connection control to obtain a plurality of replacement results, and a target replacement result is determined among the plurality of replacement results.
6. The method according to claim 1, characterized in that The method further comprises: Generate duplicate data based on business data and use it as data to be analyzed; When the preset time arrives, judging whether the business data is updated according to the data to be analyzed; If so, the data to be analyzed is updated according to the updated business data.
7. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Data processing method and device, equipment and medium
CN115576974A