Data Processing Method, Apparatus, Device, Medium, and Program Product
By executing multiple data processing instructions in data processing and writing the results into a variable table, the problem of data processing in the prior art requires a lot of manual participation, and more efficient and accurate automated data processing is achieved.
Patent Information
- Application Number
- CN202111558148.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-17
AI Technical Summary
In the prior art, data processing requires a lot of manual participation, resulting in large workload, cumbersome statistics, long time-consuming, and possible errors, resulting in inaccurate results.
By obtaining the associated data in the source file, executing multiple data processing instructions to obtain the values of statistical indicators, and writing these values into the variable table, and then replacing the variables in the preset document template according to the content in the variable table to achieve automated data processing.
It reduces manual participation, improves the real-time and accuracy of data processing, and improves the flexibility of automatic processing of scenarios, avoiding possible errors and time-consuming problems in manual statistics.
Smart Images

Figure CN114218925B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and more particularly, to a data processing method, apparatus, device, medium, and program product. Background Art
[0002] In daily life and work, various files are usually provided to record the data of each user. After filling in the data of each user into the file, manual participation is required for the collation, analysis, or summary of the data in the file. If the amount of data or the content of the data is large, it may pose problems such as heavy workload, cumbersome statistics, long time consumption for the processing personnel, and may also lead to statistical errors, resulting in poor real-time performance and inaccurate final results. Therefore, how to reduce manual participation and achieve automated data processing is an urgent problem to be solved currently. Summary of the Invention
[0003] In view of the above problems, the present disclosure provides a data processing method, apparatus, device, medium, and program product that realizes automated data processing to improve real-time performance and accuracy.
[0004] In one aspect of the embodiments of the present disclosure, a data processing method is provided, including: obtaining a source file, where the source file includes associated data of N users; based on the associated data, executing M data processing instructions to obtain values of M statistical metrics, where each of the M data processing instructions includes a processing condition for obtaining the value of the corresponding statistical metric, and N and M are integers greater than or equal to 1 respectively; according to a first correspondence, writing the values of the M statistical metrics into a variable table, where the variable table includes M first variables, and the first correspondence includes the correspondence between the values of the M statistical metrics and the M first variables; according to a second correspondence, based on the values of the M statistical metrics obtained from the variable table, replacing M second variables in a preset document template, where the second correspondence includes the correspondence between the M second variables and the M first variables.
[0005] According to an embodiment of the present disclosure, the method is implemented by a mixed programming method, specifically including: using a first programming language to obtain a first executable statement to implement the obtaining of the source file, the execution of the M data processing instructions, and the writing of the values of the M statistical metrics into the variable table; using a second programming language to obtain a second executable statement to implement the replacing of the M second variables in the preset document template based on the values of the M statistical metrics obtained from the variable table; where the first programming language is different from the second programming language.
[0006] According to an embodiment of the present disclosure, the second programming language is Visual Basic language, the preset document template is a preset Word template, the preset Word template includes preset text content, and the M second variables are set at M positions in the preset text content.
[0007] According to an embodiment of the present disclosure, the M statistical indicators include at least one classification indicator, each classification indicator corresponds to at least one keyword, and obtaining the values of the M statistical indicators includes obtaining the values of each classification indicator, specifically including: matching each keyword in the at least one keyword with a field in the associated data; accumulating the number of successful matches of each keyword with the field in the associated data to obtain the total number of matches; and using the total number of matches as the value of the corresponding classification indicator.
[0008] According to an embodiment of the present disclosure, it further includes setting the priority order of each classification indicator, and the accumulating the number of successful matches of each keyword with the field in the associated data includes: when the keywords of multiple classification indicators are respectively successfully matched with the fields at the same position in the associated data, accumulating the number of successful matches of the keyword with the highest priority; wherein, the keyword with the highest priority corresponds to the classification indicator with the highest priority among the multiple classification indicators, and the same position includes the same area in the source file.
[0009] According to an embodiment of the present disclosure, it further includes setting the priority order of each keyword in the at least one keyword, and the accumulating the number of successful matches of each keyword with the field in the associated data includes: when multiple keywords in the at least one keyword are respectively successfully matched with the fields at the same position in the associated data, accumulating the number of successful matches of the keyword with the highest priority, wherein the same position includes the same area in the source file.
[0010] According to an embodiment of the present disclosure, the associated data includes at least one kind of assessment data of the N users, the source file is an Excel file, the Excel file includes at least one sheet page, and each sheet page includes one kind of assessment data. Obtaining the source file includes obtaining the Excel file. After obtaining the Excel file, the method further includes: performing an entry operation to write the assessment data in each sheet page into a corresponding first database table; wherein, the performing M data processing instructions based on the associated data to obtain the values of the M statistical indicators includes: performing the corresponding at least one data processing instruction based on the type of assessment data in each first database table.
[0011] According to an embodiment of the present disclosure, the M statistical indicators include M assessment indicators. Before executing the M data processing instructions, it further includes presetting at least one data processing instruction corresponding to each type of assessment data, specifically including: determining at least one assessment indicator corresponding to each type of assessment data; and presetting the data processing instruction corresponding to each assessment indicator according to the processing condition of the value of each assessment indicator in the at least one assessment indicator.
[0012] Another aspect of the embodiments of the present disclosure provides a data processing device, including: a file acquisition module for acquiring a source file, where the source file includes associated data of N users; an instruction execution module for executing M data processing instructions based on the associated data to obtain the values of M statistical indicators, where each data processing instruction in the M data processing instructions includes a processing condition for obtaining the value of the corresponding statistical indicator, and N and M are integers greater than or equal to 1 respectively; an indicator writing module for writing the values of the M statistical indicators into a variable table according to a first correspondence, where the variable table includes M first variables, and the first correspondence includes the correspondence between the values of the M statistical indicators and the M first variables; a template replacement module for replacing M second variables in a preset document template based on the values of the M statistical indicators obtained from the variable table according to a second correspondence, where the second correspondence includes the correspondence between the M second variables and the M first variables.
[0013] Another aspect of the embodiments of the present disclosure provides an electronic device, including: one or more processors; a storage device for storing one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method as described above.
[0014] Another aspect of the embodiments of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the method as described above.
[0015] Another aspect of the embodiments of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method as described above is implemented.
[0016] One or more of the above embodiments have the following beneficial effects: Based on the associated data in the source file, the values of M statistical metrics are obtained by executing M data processing instructions, and the values of the M statistical metrics are filled into a variable table, and in the variable table, the values of the M statistical metrics correspond one-to-one with M first variables. Then, based on the content in the variable table, the values of the M statistical metrics are taken out to replace M second variables in a preset document template, so as to automatically obtain the final processed document. It is possible to replace the original operation of manually obtaining the values of M statistical metrics by executing M data processing instructions, and use the variable table as the connection between the source file and the preset document template. When replacing the M second variables in the preset document template, it is no longer necessary to rely on the associated data in the source file, which improves the flexibility of the automatic processing scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above content and other objects, features and advantages of the present disclosure will become clearer. In the drawings:
[0018] Figure 1 Schematically shows an application scenario diagram of the data processing method according to an embodiment of the present disclosure;
[0019] Figure 2 Schematically shows a flowchart of the data processing method according to an embodiment of the present disclosure;
[0020] Figure 3 Schematically shows a schematic diagram of a preset document template according to an embodiment of the present disclosure;
[0021] Figure 4 Schematically shows a schematic diagram of the final output document according to an embodiment of the present disclosure;
[0022] Figure 5 Schematically shows a flowchart of the data processing method according to another embodiment of the present disclosure;
[0023] Figure 6 Schematically shows a flowchart of a preset data processing instruction according to an embodiment of the present disclosure;
[0024] Figure 7 Schematically shows a flowchart of obtaining the value of each classification metric according to an embodiment of the present disclosure;
[0025] Figure 8 Schematically shows a structural block diagram of a data processing apparatus according to an embodiment of the present disclosure;
[0026] Figure 9 Schematically shows a block diagram of an electronic device suitable for implementing the data processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present disclosure.
[0028] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0030] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0031] In the technical solution of the present disclosure, the processing of data such as acquisition, collection, storage, use, processing, transmission, provision, disclosure, and application, etc., is carried out with the permission of the user, complies with the provisions of relevant laws and regulations, takes necessary confidentiality measures, and does not violate public order and good customs.
[0032] In related technologies, for example, in scenarios such as basic situation surveys of multiple users, project progress reports, audit reports, work performance assessment reports, etc., it is necessary to collect associated data of users for statistics. Taking the work performance assessment report as an example, the daily work of each person can be recorded in an online document, which is convenient for tracking and summarizing the daily work progress and generating relevant daily reports, monthly reports, etc. Often, the above work requires manual participation, with a large workload, cumbersome statistics, long time consumption, and possible errors in manual statistics, resulting in inaccurate final report files.
[0033] Embodiments of the present disclosure provide a data processing method, including: obtaining a source file, where the source file includes associated data of N users. Based on the associated data, execute M data processing instructions to obtain values of M statistical metrics, where each of the M data processing instructions includes a processing condition for obtaining the value of the corresponding statistical metric, and N and M are integers greater than or equal to 1 respectively. According to the first correspondence, write the values of the M statistical metrics into a variable table, where the variable table includes M first variables, and the first correspondence includes the correspondence between the values of the M statistical metrics and the M first variables. According to the second correspondence, based on the values of the M statistical metrics obtained from the variable table, replace M second variables in a preset document template, where the second correspondence includes the correspondence between the M second variables and the M first variables.
[0034] According to the embodiments of the present disclosure, on the one hand, it is possible to replace the original operation of manually obtaining the values of M statistical metrics by executing M data processing instructions, avoiding possible errors, calculation errors, or long time consumption in the process of manually obtaining statistical metrics. On the other hand, if directly processing the associated data in the source file and directly filling the obtained values of the statistical metrics in the preset document template, the dependence on the source file is relatively large. In the case of a large number of source files or a large amount of data in the files, the flexibility of automatic processing is poor. In fact, the values of the statistical metrics are the data required by the preset document template. Therefore, using the variable table as the connection between the source file and the preset document template can no longer rely on the associated data in the source file when replacing the M second variables in the preset document template, improving the flexibility of the automatic processing scenario.
[0035] Figure 1 An application scenario diagram of the data processing method according to the embodiments of the present disclosure is schematically shown.
[0036] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0037] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0038] The terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on.
[0039] The server 105 can be a server that provides various services. For example, it can be a background management server (only for illustration) that supports the websites browsed by users using the terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0040] According to an embodiment of the present disclosure, a user can fill in associated data on the local device or an online web page through the terminal devices 101, 102, and 103, and save the source file including the associated data locally or on the network. The server 105 can obtain the source files stored locally on the terminal devices 101, 102, and 103, or stored on the network.
[0041] It should be noted that the data processing method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the data processing device provided by the embodiments of the present disclosure can generally be set in the server 105. The data processing method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, and 103 and / or the server 105. Correspondingly, the data processing device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, and 103 and / or the server 105.
[0042] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in
[0043] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0043] Based on the Figure 1 scenario described below, the data processing method of the embodiments of the present disclosure will be described in detail through Figures 2 to 7 the following.
[0044] Figure 2 Schematically shows a flowchart of the data processing method according to an embodiment of the present disclosure. Figure 3 Schematically shows a schematic diagram of a preset document template according to an embodiment of the present disclosure. Figure 4 Schematically shows a schematic diagram of the final output document according to an embodiment of the present disclosure.
[0045] As shown in Figure 2As shown, the data processing method of this embodiment includes operations S210 to S240.
[0046] In operation S210, a source file is obtained, where the source file includes associated data of N users.
[0047] Referring to Figure 1 , the source file can be one or more files. For example, it can be a Word file or an Excel file stored locally on terminal devices 101, 102, and 103, or it can be an online file browsed by terminal devices 101, 102, and 103 through web pages, cloud disks, etc. The associated data can be the data filled in by each of the N users into the source file. For example, the work report data filled in by each user according to the day's work content. It can also be the data filled in by a dedicated person into the source file and associated with the N users. For example, the student information collected and filled in by teachers in a school, or the patient information collected and filled in by doctors in a hospital.
[0048] In operation S220, based on the associated data, M data processing instructions are executed to obtain the values of M statistical indicators, where each of the M data processing instructions includes a processing condition for obtaining the corresponding statistical indicator, and N and M are integers greater than or equal to 1 respectively.
[0049] The value of the statistical indicator can be the content itself in the associated data, such as dates like day, week, month, or year. The processing condition is to read the content of the corresponding field. It can also be processed on the basis of the associated data. For example, the statistical value obtained on the basis of the original data. The processing condition is the specific processing sequence or calculation condition. The processing process can be realized by executing data processing instructions. Data processing instructions refer to commands that a computer can recognize and execute a certain processing operation. Data processing instructions can be formed into executable program statements by using executable functions, arithmetic operators, etc. of a programmable language, such as SQL statements.
[0050] In operation S230, according to the first correspondence, the values of the M statistical indicators are written into a variable table, where the variable table includes M first variables, and the first correspondence includes the correspondence between the values of the M statistical indicators and the M first variables.
[0051] The M first variables can be M fields representing the M statistical indicators. Table 1 exemplarily shows part of the content of the variable table of the embodiment of the present disclosure, as shown below.
[0052] Table 1
[0053]
[0054]
[0055] Among them, multiple first variables can be included in the "Variable" column, such as DATE, BG1, BG2, BG3, PL1, PL2, PL3, and WORK3. Multiple values of statistical indicators written can be included in the "Zhibiao" column. Chinese content corresponding to the variables is included in the "Remark" column.
[0056] Referring to Table 1, the first correspondence can be determined by the correspondence of each row in the variable table or in the form of key-value pairs. For example, after executing an SQL statement, the result is updated into the corresponding variable. Taking BG1 as an example, the source file can be an Excel file, in which there can be a "Change" column. Any one of version change, program change, and data change can be filled in this column. The statistical indicator corresponding to BG1 can be determined by counting the number of fields in the "Change" column, such as 24. The acquisition methods of other quantity-based indicators can be similar to that of BG1 and will not be elaborated here. Taking WORK3 as another example, the Excel file can include a "Temporary Work Content" column, and each row is the data filled in by a user. If the user has temporary work content, it is filled in; if not, it is left blank. Therefore, a data processing instruction can be set to read the content of the "Temporary Work Content" column. If content is read, WORK3 is updated. If no content is read, WORK3 is set to a null value. It should be noted that, for example, variables that need to be calculated can also be set in the variable table, such as obtained by means of average calculation, judgment, product calculation, etc., and are not limited to the content in Table 1.
[0057] In operation S240, according to the second correspondence, based on the values of the M statistical indicators obtained from the variable table, M second variables in the preset document template are replaced, where the second correspondence includes the correspondence between the M second variables and the M first variables.
[0058] Referring to Figure 3 , multiple second variables are included in the preset document template, such as DOCVARIABLE DATE, DOCVARIABLEBG1, DOCVARIABLE BG2, DOCVARIABLE BG3, DOCVARIABLE PL1, DOCVARIABLE PL2, DOCVARIABLE PL3, and DOCVARIABLE WORK3. The second correspondence can be determined by some fields in the first variable and the second variable being the same.
[0059] Referring to Figure 4 , after replacing the values of the M statistical indicators for the M second variables, a final output document is obtained, thus completing the automated data processing, greatly reducing manual operations, reducing repetitive work, and improving data accuracy.
[0060] According to an embodiment of the present disclosure, on the one hand, the execution of M data processing instructions can be used to replace the original operation of manually obtaining the values of M statistical indicators, avoiding the problems of errors, calculation errors, or long time consumption that may occur during the process of manually obtaining statistical indicators. On the other hand, if the associated data in the source file is directly processed and the obtained statistical indicators are directly filled in the preset document template, the dependence on the source file is relatively large. In the case of a large number of source files or a large amount of data in the files, the flexibility of automatic processing is poor. In fact, the values of the statistical indicators are the data required by the preset document template. Therefore, using a variable table as the connection between the source file and the preset document template can, when replacing the M second variables in the preset document template, no longer depend on the associated data in the source file, improving the flexibility of the automatic processing scenario.
[0061] According to an embodiment of the present disclosure, operations S210 to S240 can be implemented by means of hybrid programming. Specifically, it includes: obtaining first executable statements using a first programming language to implement obtaining the source file, executing M data processing instructions, and writing the values of M statistical indicators into the variable table. Obtaining second executable statements using a second programming language to implement replacing the M second variables in the preset document template based on the values of M statistical indicators obtained from the variable table. Among them, the first programming language is different from the second programming language. The following is a reference Figure 5 for further illustration.
[0062] Figure 5 The flowchart of a data processing method according to another embodiment of the present disclosure is schematically shown.
[0063] As Figure 5 shown, the data processing method of this embodiment includes operations S501 to S506.
[0064] In operation S501, third executable statements are obtained using a second programming language to generate a preset document template.
[0065] According to an embodiment of the present disclosure, the second programming language is Visual Basic language (hereinafter referred to as VB language), and the preset document template is a preset Word template. Referring to Figure 3 , the preset Word template includes preset text content, and M second variables are set at M positions in the preset text content.
[0066] An executable statement can be a statement that can notify a computer to complete one or more operations after being compiled.
[0067] The third executable statements obtained through the VB language can include a macro program created through the VB language. Through the execution of the macro program and the preset text content input manually, a preset text template can be obtained.
[0068] In operation S502, execute the first executable statement to obtain the source file.
[0069] The first programming language can be languages such as C, C++, GO, JAVA, or Python. Taking the Python language as an example, multiple first executable statements can be obtained by editing using pycharm (for example only). Obtaining the source file can be achieved by reading the statement "data = pd.read_excel(r'file path')".
[0070] In operation S503, process the source file and put it into the database in the database layer, and execute M data processing instructions to obtain the values of M statistical indicators. According to the first correspondence, write the values of the M statistical indicators into the variable table. Among them, the variable table can be the second database table stored in the database. The database table is an object used to store data in the database and is a collection of structured data.
[0071] In operation S504, call the second executable statement.
[0072] Automatically calling VB code can be achieved through a calling statement written in Python, and the VB code includes the second executable statement.
[0073] In operation S505, by executing the second executable statement, according to the second correspondence, obtain the values of the M statistical indicators from the variable table.
[0074] For example, in the VB code, it includes "sql = select * from bianliang", whose function is to connect to the database to facilitate obtaining the content in the variable table. Among them, bianliang is the name of the variable table.
[0075] In operation S506, replace the M second variables in the preset document template with the values of the M statistical indicators obtained in operation S505.
[0076] In this operation, a pre-set macro program can be run. According to the correspondence between the first variable and the second variable, obtain the statistical indicators to replace the corresponding second variable.
[0077] According to an embodiment of the present disclosure, using a variable table as the connection for statement execution between the first programming language and the second programming language can still automatically complete data processing in the way of implementing hybrid programming. In addition, the way of hybrid programming can give full play to the respective advantages of the first programming language and the second programming language. For example, the Python language can improve programming efficiency in aspects such as source file processing, file storage in the database, and data statistics. The combination of the VB language and the Word template, the combination of the logical data processing in the VB code and the logical method of Python calling VB, simplifies the functions that are relatively complex to implement in Python and facilitates the generation of the final document.
[0078] According to an embodiment of the present disclosure, the associated data includes at least one assessment data of N users, the source file is an Excel file, the Excel file includes at least one sheet page, and each sheet page includes one type of assessment data. Obtaining the source file includes obtaining the Excel file. After obtaining the Excel file, the method further includes: performing a storage operation to write the assessment data in each sheet page into a corresponding first database table.
[0079] In operation S220, based on the associated data, executing M data processing instructions to obtain the values of M statistical indicators includes: based on the types of assessment data in each first database table, executing at least one corresponding data processing instruction.
[0080] Taking the example that each person fills in the work content of the day, the Excel file can include multiple sheet pages, which are respectively used to fill in the change situation, batch problem situation, or temporary work content situation. The change situation, batch problem situation, or temporary work content situation respectively represents one type of assessment data.
[0081] Since each type of assessment data is different, the corresponding data processing instructions are also different. For example, if a first database table contains change situations, then execute the data processing instructions corresponding to the change situations to obtain assessment indicators such as the number of version changes, the number of program changes, and the number of data changes. Specifically, the name of the above first database table can be defined as "BGQK" (for example only), and several SQL statements are used to read the assessment data through "BGQK" and perform processing.
[0082] Figure 6 Schematically shows a flowchart of preset data processing instructions according to an embodiment of the present disclosure.
[0083] The M statistical indicators include M assessment indicators. Before executing the M data processing instructions in operation S220, as Figure 6 shown, it further includes presetting at least one data processing instruction corresponding to each type of assessment data, specifically including operation S610 to operation S620.
[0084] In operation S610, at least one evaluation index corresponding to each type of evaluation data is determined.
[0085] Referring to Table 1 and Figure 3 , each type of evaluation data will have corresponding evaluation indexes. For example, for the change situation, the evaluation indexes include the number of version changes, the number of program changes, the number of data changes, etc. For the batch problem situation, the evaluation indexes include the total number of batch problems, the number of problems resolved fundamentally, the number of problems to be followed up later, etc. For the ad-hoc work content, the specific work content is used as the evaluation index, such as "optimize alarm monitoring, check green light scripts, etc."
[0086] In operation S620, according to the processing conditions of the values of each evaluation index among the at least one evaluation index, data processing instructions corresponding to each evaluation index are preset.
[0087] For example, the batch problem situation also corresponds to the evaluation index of the average daily batch problems in the current month. The processing condition of the value of this index is to obtain the total number of batch problems in the current month and then divide it by the number of days in the current month. Based on this, an SQL statement can be preset as the data processing instruction corresponding to the average daily batch problems in the current month.
[0088] According to the embodiments of the present disclosure, even if the associated data in the source file changes, or the statistical index changes, or the form of the output result changes, the requirements can be flexibly adapted by changing the data processing instructions, the first variable, and the preset document template, so as to achieve the effect of automated processing. For example, when it is desired to output in the form of a chart, it can also be achieved through VB code.
[0089] Figure 7 A flowchart for obtaining each classification index according to an embodiment of the present disclosure is schematically shown.
[0090] As Figure 7 shown, in operation S220, based on the associated data, executing M data processing instructions to obtain the values of M statistical indexes may include obtaining the values of each classification index, such as operations S710 to S730. Among them, the M statistical indexes include at least one classification index, and each classification index corresponds to at least one keyword.
[0091] In operation S710, each keyword among the at least one keyword is matched with a field in the associated data.
[0092] Table 2 schematically shows the keyword content configured for each classification index in the embodiments of the present disclosure, as shown below.
[0093] Table 2
[0094] Classification Index Keyword 1 Keyword 2 Keyword 3 Program Problem Program Restart Version Data Problem Duplicate Data Not Unique Database Problem Database Timeout Connection Table Structure Problem Source Table Structure Null
[0095] Referring to Table 1, taking program problems as an example, which correspond to the three keywords "program", "restart", and "version", they can be respectively matched with the fields in the associated data.
[0096] In operation S720, accumulate the number of successful matches of each keyword with the fields in the associated data to obtain the total number of matches.
[0097] For example, an Excel file includes a sheet named "Batch Problem Situation", which contains relevant data on batch problems. This sheet includes a column named "Batch Failure Reason", and each row of data corresponding to this column is the specific reason content filled in by the user, such as "When executing a certain SQL statement on a certain database table in the database, an error occurs because the table structures of the source table and the target table are inconsistent and data cannot be inserted, resulting in an error".
[0098] Taking the above reason content as an example, in the "data problem" indicator, the keyword "data" can match the "database" field and the "data" field in the above reason content, that is, the accumulated number of times is 2 times. In the "database problem" indicator, the keyword "database" can match the "database" field in the above reason content, that is, the accumulated number of times is 1 time. In the "table structure problem" indicator, the keyword "source table" can match the "source table" field in the above reason content, and the keyword "structure" can match the "table structure" field in the above reason content, that is, the accumulated number of times is 2 times. The fields in each row of data in the "Batch Failure Reason" column can be matched in a similar way as the above accumulation method to obtain the total number of matches.
[0099] In operation S730, take the total number of matches as the value of the corresponding classification indicator.
[0100] According to the embodiments of the present disclosure, by associating data content and keywords, an automated data classification function is achieved. Compared with manual classification by reading a large amount of associated data, the classification efficiency is improved, and the situation of manual classification errors is avoided.
[0101] According to the embodiments of the present disclosure, it also includes setting the priority order of each classification indicator. Accumulating the number of successful matches of each keyword with the fields in the associated data includes: when the keywords of multiple classification indicators are respectively successfully matched with the fields in the same position in the associated data, accumulating the number of successful matches of the keyword with the highest priority. Among them, the keyword with the highest priority corresponds to the classification indicator with the highest priority among multiple classification indicators, and the same position includes the same area in the source file.
[0102] The same area in the source file can be the area of the same row or the same column in the Excel file, or the same cell area located by row coordinates and column coordinates. It can also be the area of the same sentence or the same paragraph in a certain cell.
[0103] If, after troubleshooting, it is determined that the problem of "an error occurs when executing a certain SQL statement on a certain database table in the database, the table structures of the source table and the target table are inconsistent, and the data cannot be inserted, resulting in an error" is a table structure problem. Then, the accumulation of the "data problem" indicator and the "database problem" indicator is redundant, which may generate data noise and interfere with the subsequent analysis results. Therefore, the "table structure problem" indicator can be set to have the highest priority, the "database problem" indicator has the second highest priority, and the "data problem" indicator has the lowest priority. When the keywords corresponding to the three classification indicators all match a field in a specific reason content, only the number of successful matches of the keyword of the "table structure problem" indicator can be effectively accumulated.
[0104] According to an embodiment of the present disclosure, the values of each classification indicator can also be obtained in sequence according to the priority order. For example, in the case where the above "table structure problem" indicator has the highest priority, first match the keywords "source table" and "structure" with the fields in the "batch failure reason" column. After a certain keyword matches successfully, delete or mark the content in this area as read to prevent the keywords of other classification indicators from matching.
[0105] According to an embodiment of the present disclosure, it also includes setting the priority order of each keyword in at least one keyword. The accumulation of the number of successful matches of each keyword with the fields in the associated data includes: when multiple keywords in at least one keyword respectively match the fields at the same position in the associated data, accumulate the number of successful matches of the keyword with the highest priority.
[0106] If, after troubleshooting, it is determined that the problem of "an error occurs when executing a certain SQL statement on a certain database table in the database, the table structures of the source table and the target table are inconsistent, and the data cannot be inserted, resulting in an error" is a table structure problem. However, both keywords "source table" and "structure" in the "table structure problem" indicator match successfully, and the accumulated number of times is 2 times. In fact, the above specific reason content only represents one occurrence of a table structure problem, and it can be considered that redundant times have been accumulated.
[0107] Therefore, the keyword "source table" can be set to have the highest priority, and the "structure" has the second highest priority. When both keywords match a field in a specific reason content, only the number of successful matches of the "source table" can be effectively accumulated.
[0108] According to an embodiment of the present disclosure, each keyword can also be matched in sequence according to the priority order. For example, in the case where the above "source table" has the highest priority, first match the keyword "source table" with the fields in the "batch failure reason" column. After the match is successful, delete or mark the content in this area as read to prevent the keyword of "structure" from matching successfully and causing redundant accumulation problems.
[0109] It should be noted that the classification indicators, the number and content of keywords in Table 2 are only examples. The priority order among the above classification indicators and the priority order among the keywords are only examples, and any settings can be made without departing from the concept of the present disclosure.
[0110] Based on the above data processing method, the present disclosure also provides a data processing device. The following will be combined with Figure 8 to describe the device in detail.
[0111] Figure 8 The structural block diagram of the data processing device according to an embodiment of the present disclosure is schematically shown.
[0112] As Figure 8 shown, the data processing device 800 in this embodiment includes a file acquisition module 810, an instruction execution module 820, an index writing module 830, and a template replacement module 840.
[0113] The file acquisition module 810 can perform operation S210 to obtain a source file, where the source file includes associated data of N users.
[0114] The instruction execution module 820 can perform operation S220 to execute M data processing instructions based on the associated data to obtain the values of M statistical indicators. Each of the M data processing instructions includes a processing condition for obtaining the corresponding statistical indicator value, and N and M are integers greater than or equal to 1 respectively.
[0115] The index writing module 830 can perform operation S230 to write the values of M statistical indicators into a variable table according to a first correspondence relationship, where the variable table includes M first variables, and the first correspondence relationship includes the correspondence relationship between the values of M statistical indicators and M first variables.
[0116] The template replacement module 840 can perform operation S240 to replace M second variables in a preset document template based on the values of M statistical indicators obtained from the variable table according to a second correspondence relationship, where the second correspondence relationship includes the correspondence relationship between M second variables and M first variables.
[0117] The data processing device 800 may further include an instruction preset module for performing operations S610 to S620. The data processing device 800 may further include a classification index calculation module for performing operations S710 to S730, which will not be elaborated here.
[0118] According to an embodiment of the present disclosure, taking assessment data as an example, the data processing device 800 can crawl assessment data such as changes, batch problems, application groups, transaction volumes, etc. from an online document, import them into a local database, and then write them into a variable table after analysis based on the content in the database, automatically generating detailed daily and monthly reports. It can also count the daily workload, weekly workload, etc. of each person based on the data in the database. And it can automatically analyze and generate daily and monthly reports according to the problem situation and change situation, saving the complexity of manual analysis, automating repetitive work, and visualizing the data.
[0119] According to an embodiment of the present disclosure, any multiple of the file acquisition module 810, the instruction execution module 820, the metric writing module 830, and the template replacement module 840 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module.
[0120] According to an embodiment of the present disclosure, at least one of the file acquisition module 810, the instruction execution module 820, the metric writing module 830, and the template replacement module 840 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Or, at least one of the file acquisition module 810, the instruction execution module 820, the metric writing module 830, and the template replacement module 840 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0121] Figure 9 A block diagram of an electronic device suitable for implementing a data processing method according to an embodiment of the present disclosure is schematically shown.
[0122] As Figure 9As shown, an electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. The processor 901 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 901 can also include on-board memory for caching purposes. The processor 901 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0123] In the RAM 903, various programs and data required for the operation of the electronic device 900 are stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to an embodiment of the present disclosure by executing the program in the ROM 902 and / or the RAM 903. It should be noted that the program can also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 can also perform various operations of the method flow according to an embodiment of the present disclosure by executing the program stored in one or more memories.
[0124] According to an embodiment of the present disclosure, the electronic device 900 can further include an input / output (I / O) interface 905, and the input / output (I / O) interface 905 is also connected to the bus 904. The electronic device 900 can further include one or more of the following components connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc. An output section 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc. A storage section 908 including a hard disk, etc. And a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read from it can be installed into the storage section 908 as needed.
[0125] The present disclosure also provides a computer-readable storage medium, which can be included in the device / device / system described in the above embodiment. It can also exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to an embodiment of the present disclosure is implemented.
[0126] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the above-described ROM 902 and / or RAM 903 and / or ROM 902 and RAM 903.
[0127] An embodiment of the present disclosure also includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method provided by the above embodiments of the present disclosure.
[0128] When the computer program is executed by the processor 901, it executes the above functions defined in the system / apparatus of the embodiments of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0129] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 909, and / or be installed from the removable medium 911. The program code included in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0130] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or be installed from the removable medium 911. When the computer program is executed by the processor 901, it executes the above functions defined in the system of the embodiments of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0131] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0133] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0134] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A data processing method, comprising: obtaining a source file, where the source file includes associated data of N users, and the source file includes an online file or a file stored locally on a terminal device; executing M data processing instructions based on the associated data to obtain values of M statistical metrics, where each of the M data processing instructions includes a processing condition for obtaining the value of a corresponding statistical metric, and N and M are integers greater than or equal to 1 respectively; writing the values of the M statistical metrics into a variable table located in a database according to a first correspondence, where the variable table includes M first variables, and the first correspondence includes the correspondence between the values of the M statistical metrics and the M first variables; replacing M second variables in a preset document template based on the values of the M statistical metrics obtained from the variable table according to a second correspondence to obtain a final output document, where the second correspondence includes the correspondence between the M second variables and the M first variables; wherein, the method is implemented by a hybrid programming method, specifically including: executing first executable statements obtained using a first programming language to implement the obtaining of the source file, the execution of the M data processing instructions, and the writing of the values of the M statistical metrics into the variable table; after writing the values of the M statistical metrics into the variable table, invoking second executable statements obtained using a second programming language through a call statement written in the first programming language; executing the second executable statements to implement replacing the M second variables in the preset document template based on the values of the M statistical metrics obtained from the variable table; wherein, the first programming language is different from the second programming language.
2. The method according to claim 1, wherein, the second programming language is Visual Basic language, the preset document template is a preset Word template, the preset Word template includes preset text content, and the M second variables are set at M positions in the preset text content.
3. The method according to claim 1, wherein, the M statistical metrics include at least one classification metric, and each classification metric corresponds to at least one keyword. Obtaining the values of the M statistical metrics includes obtaining the values of each classification metric, specifically including: matching each keyword in the at least one keyword with a field in the associated data; accumulating the number of successful matches between each keyword and the field in the associated data to obtain a total number of matches; using the total number of matches as the value of the corresponding classification metric.
4. The method according to claim 3, wherein, it further includes setting a priority order for each classification metric, and the accumulating the number of successful matches between each keyword and the field in the associated data includes: when keywords of multiple classification metrics are successfully matched with fields at the same position in the associated data, accumulating the number of successful matches of the keyword with the highest priority. Among them, the keyword with the highest priority corresponds to the classification index with the highest priority among the multiple classification indexes, and the same position includes the same area in the source file.
5. The method according to claim 4, wherein, it further includes setting the priority order of each keyword in the at least one keyword, and the cumulative number of times each keyword matches successfully with the fields in the associated data includes: when multiple keywords in the at least one keyword respectively match successfully with the fields in the same position in the associated data, the cumulative number of times the keyword with the highest priority matches successfully, wherein the same position includes the same area in the source file.
6. The method according to claim 1, wherein, the associated data includes at least one kind of assessment data of the N users, the source file is an Excel file, the Excel file includes at least one sheet page, each sheet page includes one kind of assessment data, the obtaining of the source file includes obtaining the Excel file, and after obtaining the Excel file, the method further includes: performing a warehousing operation to write the assessment data in each sheet page into a corresponding first database table; wherein, the performing of M data processing instructions based on the associated data to obtain the values of M statistical indexes includes: performing corresponding at least one data processing instruction based on the type of assessment data in each first database table.
7. The method according to claim 6, wherein, the M statistical indexes include M assessment indexes, and before performing the M data processing instructions, it further includes presetting at least one data processing instruction corresponding to each kind of assessment data, specifically including: determining at least one assessment index corresponding to each kind of assessment data; presetting the data processing instruction corresponding to each assessment index according to the processing condition of the value of each assessment index in the at least one assessment index.
8. A data processing device, comprising: a file acquisition module, configured to acquire a source file, wherein the source file includes associated data of N users, and the source file includes an online file or a file stored locally in a terminal device; an instruction execution module, configured to perform M data processing instructions based on the associated data to obtain the values of M statistical indexes, wherein each data processing instruction in the M data processing instructions includes a processing condition for obtaining the value of the corresponding statistical index, and N and M are integers greater than or equal to 1 respectively; an index writing module, configured to write the values of the M statistical indexes into a variable table located in a database according to a first correspondence relationship, wherein the variable table includes M first variables, and the first correspondence relationship includes the correspondence relationship between the values of the M statistical indexes and the M first variables; a template replacement module, configured to replace M second variables in a preset document template based on the values of the M statistical indexes obtained from the variable table according to a second correspondence relationship to obtain a final output document, wherein the second correspondence relationship includes the correspondence relationship between the M second variables and the M first variables; Among them, the following operations are implemented through hybrid programming: Execute the first executable statement obtained using the first programming language to implement the obtaining of the source file, the execution of the M data processing instructions, and the writing of the values of the M statistical metrics into the variable table; After writing the values of the M statistical metrics into the variable table, call the second executable statement obtained using the second programming language through a call statement written in the first programming language; Execute the second executable statement to implement the replacement of M second variables in the preset document template based on the values of the M statistical metrics obtained from the variable table; Among them, the first programming language is different from the second programming language.
9. An electronic device, including: One or more processors; A storage device for storing one or more programs, wherein, when the one or more programs are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor executes the method according to any one of claims 1 to 7.
11. A computer program product, including a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Mail generation method and device, computer apparatus and storage medium
CN109523236A