Data processing method, device and equipment
By adjusting the arrangement of tabular data and performing sequence summary extraction processing, the problem of low efficiency and accuracy of tabular data analysis is solved, and more efficient business processing is achieved.
Patent Information
- Application Number
- CN202510269185.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art has low efficiency and accuracy in tabular data analysis, affecting business processing efficiency.
By adjusting the arrangement of table data to a preset arrangement method and performing sequence summary extraction processing, the similarity between table data is determined and analysis efficiency and accuracy are improved.
It improves the efficiency and accuracy of table data analysis, reduces time complexity, and improves the efficiency of subsequent business processing.
Smart Images

Figure CN120278507A_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of computer technology, and particularly to a data processing method, apparatus, and device. Background Art
[0002] With the rapid development of the Internet industry, the amount of data generated during the business service provision process is increasing. Tabular data has become one of the most widely distributed data types. Therefore, how to improve the business processing efficiency by enhancing the tabular data analysis efficiency to better provide business services for users (such as in risk detection processing, how to improve the risk detection efficiency by enhancing the tabular data analysis efficiency to protect the privacy data of users from being leaked, etc.) has become the focus of attention of network operators.
[0003] When analyzing tabular data, the intersection and union of the same fields between any two tabular data can be calculated to determine the tabular analysis result between any two tabular data, and then the corresponding business can be executed according to the tabular analysis result. However, due to the increasing number and types of fields included in tabular data, the analysis efficiency and accuracy of tabular analysis by the above method are low, affecting the subsequent business processing efficiency. Therefore, the embodiments of this specification provide a technical solution to improve the business processing efficiency by enhancing the tabular data analysis efficiency. Summary of the Invention
[0004] The purpose of the embodiments of this specification is to provide a technical solution to improve the business processing efficiency by enhancing the tabular data analysis efficiency.
[0005] To achieve the above technical solution, the embodiments of this specification are implemented as follows: A data processing method provided by an embodiment of this specification, the method includes: receiving a table analysis request for a target service; in response to the table analysis request, determining first table data and second table data corresponding to the table analysis request, and adjusting the arrangement manner of the table content of the first table data and the second table data to a preset arrangement manner; according to the adjusted arrangement manner of the first table data and the second table data, determining the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the first table data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the second table data; based on the data included in each sequence in the determined first table data, performing a summary extraction process on each sequence in the first table data to obtain summary data corresponding to each sequence in the first table data, and based on the data included in each sequence in the determined second table data, performing a summary extraction process on each sequence in the second table data to obtain summary data corresponding to each sequence in the second table data; according to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data, determining the similarity between the first table data and the second table data; according to the similarity between the first table data and the second table data, determining a table analysis result, and according to the table analysis result, executing the target service.
[0006] A data processing device provided by an embodiment of this specification, the device includes: a request receiving module, configured to receive a form analysis request for a target service; a form determining module, configured to, in response to the form analysis request, determine first form data and second form data corresponding to the form analysis request, and adjust the arrangement manner of the form content of the first form data and the second form data to a preset arrangement manner; a data determining module, configured to determine, according to the adjusted arrangement manner of the first form data and the second form data, the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the first form data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the second form data; an abstract extraction module, configured to perform an abstract extraction process on each sequence in the first form data based on the data included in each sequence determined in the first form data, to obtain abstract data corresponding to each sequence in the first form data, and perform an abstract extraction process on each sequence in the second form data based on the data included in each sequence determined in the second form data, to obtain abstract data corresponding to each sequence in the second form data; a form matching module, configured to determine the similarity between the first form data and the second form data according to the abstract data corresponding to each sequence in the first form data and the abstract data corresponding to each sequence in the second form data; a service execution module, configured to determine a form analysis result according to the similarity between the first form data and the second form data, and execute the target service according to the form analysis result.
[0007] A data processing device provided by an embodiment of this specification, the data processing device includes: a processor; and a memory arranged to store computer-executable instructions, the executable instructions, when executed, cause the processor to receive a table analysis request for a target service; in response to the table analysis request, determine first table data and second table data corresponding to the table analysis request, and adjust the arrangement manner of the table contents of the first table data and the second table data to a preset arrangement manner; according to the adjusted arrangement manner of the first table data and the second table data, determine the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the first table data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the second table data; based on the data included in each sequence determined in the first table data, perform a summary extraction process on each sequence in the first table data to obtain summary data corresponding to each sequence in the first table data, and based on the data included in each sequence determined in the second table data, perform a summary extraction process on each sequence in the second table data to obtain summary data corresponding to each sequence in the second table data; according to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data, determine the similarity between the first table data and the second table data; according to the similarity between the first table data and the second table data, determine a table analysis result, and execute the target service according to the table analysis result.
[0008] An embodiment of this specification also provides a storage medium, which is used to store computer-executable instructions. When the executable instructions are executed by a processor, the following processes are implemented: receiving a table analysis request for a target service; in response to the table analysis request, determining first table data and second table data corresponding to the table analysis request, and adjusting the arrangement mode of the table content of the first table data and the second table data to a preset arrangement mode; according to the adjusted arrangement modes of the first table data and the second table data, determining the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement mode in the first table data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement mode in the second table data; based on the data included in each sequence determined in the first table data, performing a summary extraction process on each sequence in the first table data to obtain summary data corresponding to each sequence in the first table data, and based on the data included in each sequence determined in the second table data, performing a summary extraction process on each sequence in the second table data to obtain summary data corresponding to each sequence in the second table data; determining the similarity between the first table data and the second table data according to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data; determining a table analysis result according to the similarity between the first table data and the second table data, and executing the target service according to the table analysis result. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings. Figure 1 This is an embodiment of a data processing method in this specification; Figure 2 This is a schematic diagram of web page data in this specification; Figure 3 This is another embodiment of a data processing method in this specification; Figure 4 This is another embodiment of a data processing method in this specification; Figure 5 This is another embodiment of a data processing method in this specification; Figure 6 This is another embodiment of a data processing method in this specification; Figure 7This is another embodiment of the data processing method described in this specification; Figure 8 This is another embodiment of the data processing method described in this specification; Figure 9 This is another embodiment of the data processing method described in this specification; Figure 10 This is an embodiment of a data processing device described in this specification; Figure 11 This is an embodiment of a data processing device described in this specification. Detailed implementation manners
[0010] The embodiments of this specification provide a data processing method, device, and equipment.
[0011] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0012] The embodiments of this specification provide a processing mechanism for improving business processing efficiency by enhancing table data analysis efficiency. With the rapid development of the Internet industry, the amount of data generated during business service provision is increasing, and tabular data has become one of the most widely distributed data types. Therefore, how to improve business processing efficiency by enhancing tabular data analysis efficiency to better provide business services to users (such as how to improve risk detection efficiency by enhancing tabular data analysis efficiency to protect users' private data from being leaked during risk detection processing, etc.) has become the focus of attention for network operators. When analyzing tabular data, the intersection and union of the same fields between any two tabular data can be calculated to determine the tabular analysis result between any two tabular data, and then the corresponding business can be executed based on the tabular analysis result. However, due to the increasing number and types of fields contained in tabular data, the analysis efficiency and accuracy of tabular analysis using the above method are low, affecting the subsequent business processing efficiency. For this reason, the embodiments of this specification provide a technical solution for improving business processing efficiency by enhancing tabular data analysis efficiency. In this solution, by receiving a tabular analysis request for a target business, in response to the tabular analysis request, the first tabular data and the second tabular data corresponding to the tabular analysis request are determined, and the arrangement of the tabular content of the first tabular data and the second tabular data is adjusted to a preset arrangement. According to the adjusted arrangement of the first tabular data and the second tabular data, the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement of the first tabular data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement of the second tabular data are determined. Based on the data included in each sequence of the determined first tabular data, a summary extraction process is performed on each sequence of the first tabular data to obtain the summary data corresponding to each sequence of the first tabular data, and based on the data included in each sequence of the determined second tabular data, a summary extraction process is performed on each sequence of the second tabular data to obtain the summary data corresponding to each sequence of the second tabular data. According to the summary data corresponding to each sequence of the first tabular data and the summary data corresponding to each sequence of the second tabular data, the similarity between the first tabular data and the second tabular data is determined. According to the similarity between the first tabular data and the second tabular data, the tabular analysis result is determined, and based on the tabular analysis result, the target business is executed.In this way, first, the arrangement of the table content of the table data to be analyzed (i.e., the first table data and the second table data) can be adjusted to a preset arrangement to improve the processing efficiency of subsequent table analysis. Second, for each table data to be analyzed, summary extraction processing is performed in units of sequences. When the amount of data contained in the table data to be analyzed is large, the key information contained in each sequence can be retained while compressing the data volume through the summary data corresponding to each sequence. Furthermore, through the summary data corresponding to each sequence, the similarity between the two table data to be analyzed can be quickly and accurately determined, reducing the time complexity of table analysis, improving the efficiency and accuracy of table analysis, and improving the efficiency of subsequent business processing. For specific processing, refer to the specific content in the following embodiments.
[0013] As Figure 1 shown, an embodiment of this specification provides a data processing method. The execution subject of this method can be a server. The server can be an independent server or a server cluster composed of multiple servers, etc. In this embodiment, the server is used as an example of the execution subject for detailed description. The method can specifically include the following steps: In step S102, a table analysis request for a target service is received.
[0014] Among them, the target service can be any service that needs to analyze table data. For example, the target service can be services such as abnormal data detection and redundant data cleaning, that is, the target service can be a service that detects whether there are abnormal table data, redundant table data, etc. in multiple table data.
[0015] In implementation, the server can trigger a table analysis request for the target service when receiving a user's trigger to execute the target service. Or, the server can also trigger a table analysis request for the target service when reaching a preset table analysis period (such as three days, one week, one month, etc.).
[0016] For example, taking the target service as the redundant data cleaning service as an example, the user can trigger the execution of the redundant data cleaning service for a certain database. At this time, the server can trigger a table analysis request for multiple table data contained in the database. Or, the server can trigger the redundant data cleaning service for a certain database every three days (i.e., the preset table analysis period), that is, the server can trigger a table analysis request for multiple table data contained in the database every three days.
[0017] In step S104, in response to the table analysis request, the first table data and the second table data corresponding to the table analysis request are determined, and the arrangement of the table content of the first table data and the second table data is adjusted to a preset arrangement.
[0018] Among them, the first table data and the second table data can be table data determined according to the target service. For example, taking the target service as the redundant data cleaning service for a certain database, the first table data and the second table data can be any two table data in the database. Or, taking the target service as the abnormal data detection, the first table data and the second table data can include table data without abnormal data and table data that may have abnormal data, etc.
[0019] In implementation, when receiving a table analysis request for the target service, the server can first obtain the target data corresponding to the target service, and then determine the first table data and the second table data according to the target data.
[0020] Among them, the target data can be any type of data containing table data. For example, the target data can be picture data, text data, web page data, etc. of any type containing table data. The arrangement of the table content can include various ways such as horizontal arrangement, vertical arrangement, and custom arrangement. The preset arrangement way can be horizontal arrangement or vertical arrangement.
[0021] Since there are format differences in data of different file types, table data extraction processing can be performed on different types of target data to perform table analysis processing based on the extracted table data.
[0022] For example, the target service can be an abnormal data detection service. The target data corresponding to this target service can include picture data 1, text data 1, and web page data 1. Among them, the web page data 1 does not contain abnormal data, and the text data 1 and the picture data 1 may contain abnormal data. The server can perform table data extraction processing on each target data respectively, and determine the first table data and the second table data according to the table data extraction results.
[0023] Specifically, the server can determine the first table data according to the result obtained by performing table data extraction processing through network element data 1, determine the second table data 1 according to the result obtained by performing table data extraction processing through text data 1, and determine the second table data 2 according to the result obtained by performing table data extraction processing through picture data 1. Then, perform table analysis processing on the first table data and the second table data 1, and the first table data and the second table data 2 respectively.
[0024] Among them, for different types of target data, the server can adopt different methods for extracting and processing tabular data. For example, for image data, the server can perform structured analysis on the image data to extract the tabular data in the image data. The structured analysis can include two aspects: one is to extract the data structure information in the image data, and the other is to extract the text content in the image data. Then, based on the data structure information and the text content, tabular data corresponding to the image data is generated. For text data, the server can perform semantic parsing processing on the text data to determine the tabular data corresponding to the text data according to the semantic parsing result. For web page data, the server can perform code parsing processing on the source file of the network element data to determine the tabular data corresponding to the web page data.
[0025] In this way, through different methods for extracting tabular data, it is possible to adapt to tabular data of multiple file carriers and achieve similarity calculation between cross-file type tables.
[0026] In addition, the above method for determining the first tabular data and the second tabular data is an optional and implementable determination method. In actual application scenarios, there can be various different determination methods, and different determination methods can be selected according to different actual application scenarios. This specification does not make specific limitations on this.
[0027] The server can respectively determine the arrangement methods of the tabular contents of the first tabular data and the second tabular data, and in the case where the arrangement methods of the tabular contents of the first tabular data and / or the second tabular data are not the preset arrangement methods, perform adjustment processing on the arrangement methods to adjust the arrangement methods of the tabular contents of the first tabular data and the second tabular data to the preset arrangement methods.
[0028] For example, taking the preset arrangement method as vertical arrangement as an example, assuming that the above first tabular data is the tabular data determined by the server according to the horizontal arrangement method for the code parsing result obtained by performing code parsing processing on the web page data 1 shown as follows Figure 2 Specifically, the generated first tabular data can be as shown in Table 1 below.
[0029] Table 1
[0030] Since the preset arrangement method is vertical arrangement, therefore, the arrangement method of the tabular content of the first tabular data in horizontal arrangement can be adjusted to vertical arrangement. For example, the first tabular data can be transposed to obtain the adjusted arrangement method as shown in Table 2 below.
[0031] Table 2
[0032] The above takes the case where the preset arrangement is a vertical arrangement as an example. In actual application scenarios, there can be multiple preset arrangements. According to different actual application scenarios, different preset arrangements can be selected to adjust the arrangement of the table content of the first table data and the second table data. This specification embodiment does not make specific limitations on this.
[0033] In step S106, according to the adjusted arrangement of the first table data and the second table data, determine the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement of the first table data and the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement of the second table data.
[0034] In implementation, for example, taking the adjusted first table data as the table data shown in Table 2 above as an example, the corresponding arrangement direction of the adjusted arrangement in the first table data can include four sequences, namely sequence 1 (time, 1-1, 1-2, 1-3), sequence 2 (location, location 1, location 2, location 3), sequence 3 (user representation, user 1, user 1, user 2), and sequence 4 (resource transfer quantity, 100, 200, 300).
[0035] In step S108, based on the data included in each sequence of the determined first table data, perform summary extraction processing on each sequence in the first table data to obtain the summary data corresponding to each sequence in the first table data, and based on the data included in each sequence of the determined second table data, perform summary extraction processing on each sequence in the second table data to obtain the summary data corresponding to each sequence in the second table data.
[0036] In implementation, the server can use a pre-trained summary extraction model to perform summary extraction processing on each sequence based on the data included in each sequence to obtain the summary data corresponding to each sequence. Among them, the preset summary extraction model can be a model constructed based on a preset deep learning algorithm.
[0037] Alternatively, the server can also determine the summary data corresponding to each sequence according to the data type, data quantity, maximum value, minimum value, median, high-frequency value, etc. in the data included in each sequence.
[0038] The above method for summary extraction processing is an optional and implementable processing method. In actual application scenarios, there can be multiple different processing methods. According to different actual application scenarios, different processing methods can be selected. This specification embodiment does not make specific limitations on this.
[0039] In addition, before performing the abstract extraction process, the data included in each sequence can also be preprocessed. For example, the preprocessing can include data cleaning processes (such as missing value detection, missing value supplementation, outlier detection, and outlier handling, etc.), and data type recognition processes (such as identifying the data type corresponding to each sequence, where the data type can include numeric type, date type, text type, and currency type, etc.).
[0040] In step S110, according to the abstract data corresponding to each sequence in the first table data and the abstract data corresponding to each sequence in the second table data, determine the similarity between the first table data and the second table data.
[0041] In implementation, the server can use a pre-trained similarity determination model to determine the similarity between the first table data and the second table data according to the abstract data corresponding to each sequence in the first table data and the abstract data corresponding to each sequence in the second table data. Among them, the similarity determination model can be a model constructed based on a preset machine learning algorithm.
[0042] Alternatively, the server can also determine the similarity between any sequence in the first table data and any sequence in the first table data according to the abstract data corresponding to each sequence in the first table data and the abstract data corresponding to each sequence in the second table data, and then determine the similarity between the first table data and the second table data according to the similarity between any sequence in the first table data and any sequence in the first table data.
[0043] In addition, the method for determining the similarity between the above-mentioned first table data and the second table data is an optional and implementable determination method. In actual application scenarios, there can also be various different determination methods, and different determination methods can be selected according to the different actual application scenarios. This specification embodiment does not make specific limitations on this.
[0044] In step S112, according to the similarity between the first table data and the second table data, determine the table analysis result, and execute the target business according to the table analysis result.
[0045] In implementation, taking the target business as the redundant data cleaning business as an example, the server can determine whether there is redundant table data (i.e., the table analysis result) in the preset database according to the similarity between the first table data and the second table data. If there is redundant table data, the redundant data cleaning business for this database can be triggered. For example, the table data to be cleaned can be determined according to the similarity between the first table data and the second table data, and the redundant data cleaning process can be performed according to the table data to be cleaned.
[0046] Alternatively, taking the target business as an example of abnormal data detection, assume that the first table data is table data without abnormal data. The server can determine whether there is abnormal data in the second table data (i.e., the table analysis result) based on the similarity between the first table data and the second table data. If it is determined that there is abnormal data in the second table data, the server can determine the sequences with abnormal data in the second table data based on the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data, so as to perform abnormal data detection processing based on the data included in the sequences with abnormal data, etc.
[0047] In this way, through the lightweight sequence summary extraction algorithm, the data volume can be compressed while retaining the key information of the table data. At the same time, it can be quickly and flexibly deployed on various platforms and systems without relying on model training. In addition, based on the sequence summary extraction algorithm, the analysis accuracy of the similarity between table data can be improved.
[0048] An embodiment of this specification provides a data processing method. By receiving a table analysis request for a target service, in response to the table analysis request, first table data and second table data corresponding to the table analysis request are determined. The arrangement of the table content of the first table data and the second table data is adjusted to a preset arrangement. According to the adjusted arrangement of the first table data and the second table data, the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement in the first table data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement in the second table data are determined. Based on the data included in each sequence in the determined first table data, a summary extraction process is performed on each sequence in the first table data to obtain summary data corresponding to each sequence in the first table data. And based on the data included in each sequence in the determined second table data, a summary extraction process is performed on each sequence in the second table data to obtain summary data corresponding to each sequence in the second table data. According to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data, the similarity between the first table data and the second table data is determined. According to the similarity between the first table data and the second table data, a table analysis result is determined. And according to the table analysis result, the target service is executed. In this way, first, the arrangement of the table content of the table data to be analyzed (i.e., the first table data and the second table data) can be adjusted to a preset arrangement to improve the processing efficiency of subsequent table analysis. Second, for each table data to be analyzed, a summary extraction process is performed in units of sequences. When the amount of data included in the table data to be analyzed is large, while retaining the key information included in each sequence through the summary data corresponding to each sequence, the data volume can be compressed. Furthermore, through the summary data corresponding to each sequence, the similarity between the two table data to be analyzed can be quickly and accurately determined, reducing the time complexity of table analysis, improving the efficiency and accuracy of table analysis, and improving the efficiency of subsequent service processing.
[0049] In practical applications, the arrangement of the table content of the first table data is determined according to the data type, data length, and data frequency of the data in the target row in the first table data, and the data type, data length, and data frequency of the data in the target column in the first table data, where the target row can be any row in the first table data, and the target column can be any column in the first table data.
[0050] Since the data types of the same kind of data are the same (e.g., mobile phone numbers can be data of the numeric type, names can be data of the string type, etc.), the data length distributions of the same kind of data are concentrated, and the occurrence frequencies of a certain data in the same kind of data are similar. Therefore, the arrangement method of the table content of the first table data can be determined according to the data type, data length, and data frequency of the data in the target row and the target column.
[0051] For example, if the data types of the data in the target row are the same, the maximum value of the difference in the data lengths of the data is less than the preset difference threshold, and the data frequency of a certain data in the data is greater than the preset frequency value, and, the data types of the data in the target column are different, the maximum value of the difference in the data lengths of the data is greater than the preset difference threshold, and there is no data frequency of a certain data in the data that is greater than the preset frequency value. Therefore, it can be determined that the arrangement method of the table content of the first table data is horizontal arrangement. Conversely, the arrangement method of the table content of the first table data is vertical arrangement.
[0052] In this way, according to the performance of the row data and column data in the table data in the three dimensions of data type, data length, and data frequency, the arrangement method of the table content of the table data can be quickly and accurately determined, which can improve the analysis efficiency of subsequent table analysis.
[0053] In practical applications, the specific processing method for determining the similarity between the first table data and the second table data according to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data can be various. The following provides an optional processing method, such as Figure 3 shown, which can specifically include the processing of the following steps S1102~S1106.
[0054] In step S1102, according to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data, the sequence similarity between any two sequences in the first table data and the second table data is determined.
[0055] In practical applications, the summary data includes multiple feature data. The specific processing method for determining the sequence similarity between any two sequences in the first table data and the second table data according to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data in step S1102 can be various. The following provides an optional processing method, which can specifically include the processing of the following steps A1~A2.
[0056] In step A1, according to each feature data in the summary data corresponding to each sequence in the first table data and each feature data in the summary data corresponding to each sequence in the second table data, determine the similarity between any two sequences in the first table data and the second table data for each feature data.
[0057] In implementation, taking the feature data including data type, data quantity, maximum value, minimum value, median, and high-frequency value as an example, for the feature data of data type, the server can determine whether the data types in the summary data of the two sequences are the same to determine the similarity of the two sequences for the feature of data type. For example, if the data types in the summary data of the two sequences are the same, it can be determined that the similarity of the two sequences for the feature of data type is 1, and if the data types in the summary data of the two sequences are different, it can be determined that the similarity of the two sequences for the feature of data type is 0.
[0058] For the feature data of data quantity, the server can determine the ratio between the data quantities in the summary data of the two sequences as the similarity of the two sequences for the feature of data quantity. Among them, if there is a situation where the data quantity is 0 in the two sequences, the similarity of the two sequences for the feature of data quantity is 0.
[0059] For the three feature data of maximum value, minimum value, and median, the server determines the ratio between the relevant feature data in the summary data of the two sequences as the similarity of the two sequences for the corresponding feature data. Among them, if there is a situation where the feature data is 0 in the two sequences, the similarity of the two sequences for the feature data is 0.
[0060] For the feature data of high-frequency value, the server can determine the overlap degree (Intersection over Union, IoU) between the high-frequency values as the similarity of the two sequences for the feature of high-frequency value.
[0061] In step A2, according to the preset weight value corresponding to each feature data and the similarity between any two sequences in the first table data and the second table data for each feature data, determine the sequence similarity between any two sequences in the first table data and the second table data.
[0062] In implementation, the server can perform weighted processing according to the preset weight value corresponding to each feature data and the similarity between any two sequences in the first table data and the second table data for each feature data to obtain the sequence similarity between any two sequences in the first table data and the second table data.
[0063] In practical applications, before the above step A2, the preset weight value corresponding to each feature data can be determined. The method for determining the preset weight value can be various. The following provides an optional processing method, which can specifically include the processing steps B1 to B4 as follows.
[0064] In step B1, obtain the sample sequence data corresponding to the target service, the first sequence data having a preset similarity relationship with the sample sequence data, and the second sequence data having no preset similarity relationship with the sample sequence data.
[0065] Among them, the sample sequence data can be any one of the sequence data in any first historical table data. The first sequence data can be one or more sequence data in the second historical table data that have a preset correlation relationship with the sample sequence data. The first sequence data can be one or more sequence data in the second historical table data that have no preset correlation relationship with the sample sequence data. The first historical table data and the second historical table data can be table data constructed based on the historical service data corresponding to the target service. For example, the sample sequence data can be the sequence data with the sequence header feature of "name" in the first historical table data. The first sequence data 1 can be the sequence data with the sequence header feature of "name" in the second historical table data 1. The first sequence data 2 can be the sequence data with the sequence header feature of "name" in the second historical table data 2. The second sequence data 1 can be the sequence data with the sequence header features of "department" respectively in the second historical table data 1. The second sequence data 2 can be the sequence data with the sequence header features of "position" respectively in the second historical table data 1.
[0066] In step B2, according to the summary data corresponding to the sample sequence data, the summary data corresponding to the first sequence data, the summary data corresponding to the second sequence data, and the initial weight value corresponding to each feature data, respectively determine the first similarity between the sample sequence data and each first sequence data, and the second similarity between the sample feature data and each second sample data.
[0067] In implementation, the server can perform weighted calculation according to the initial weight value corresponding to each feature data to respectively determine the first similarity between the sample sequence data and each first sequence data, and the second similarity between the sample feature data and each second sample data.
[0068] In step B3, according to the first similarity and the second similarity, determine whether the initial weight value meets the preset screening requirements. If the initial weight value meets the preset screening requirements, determine the initial weight value as the preset weight value corresponding to each feature data.
[0069] In implementation, when the first similarity is greater than a preset first similarity threshold and the second similarity is less than a preset second similarity threshold, it can be determined that the initial weight value meets the preset screening requirements, where the first similarity threshold is greater than the second similarity threshold.
[0070] In step B4, when the initial weight value does not meet the preset screening requirements, the initial weight value is adjusted, and based on the adjusted initial weight value, the first similarity and the second similarity are continuously determined until the adjusted initial weight value meets the preset screening requirements, and the adjusted initial weight value is determined as the preset weight value corresponding to each feature data.
[0071] In step S1104, based on the sequence similarity between any two sequences in the first table data and the second table data, sequence matching processing is performed on the first table data and the second table data to determine the matching sequence pairs in the first table data and the second table data.
[0072] In implementation, for a sequence in the first table data, first, a sequence in the second table data that matches it needs to be found, and then based on the abstract data of the two matching sequences, the similarity between the two sequences is determined. The server can use a pre-trained matching model to perform sequence matching processing on the first table data and the second table data based on the sequence similarity between any two sequences in the first table data and the second table data, and determine the matching sequence pairs in the first table data and the second table data, where the matching model can be a model constructed according to a preset machine learning algorithm.
[0073] Alternatively, when determining the matching sequence pairs, this problem can be regarded as an assignment problem: it is necessary to pair the sequences of the two table data pairwise so that the sum of the similarities of the paired sequences is the highest. Therefore, the server can also use the Hungarian algorithm to construct a similarity matrix based on the negative values of the sequence similarities between any two sequences in the first table data and the second table data, and perform sequence matching processing on the first table data and the second table data according to the similarity matrix to determine the matching sequence pairs in the first table data and the second table data.
[0074] Among them, the Hungarian algorithm is a combinatorial optimization algorithm that can be used to find the task assignment scheme with the minimum cost. Only by replacing the cost in the Hungarian algorithm with the negative value of the similarity (the goal is to maximize the similarity, that is, to minimize the negative value of the similarity) can the matching problem between sequences be solved by the Hungarian algorithm.
[0075] In addition, the method for determining the above sequence pairs is an optional and implementable determination method. In actual application scenarios, there can be multiple different determination methods, and different determination methods can be selected according to different actual application scenarios. The embodiments of this specification do not make specific limitations on this.
[0076] In step S1106, according to the similarity between the two sequences in the matched sequence pair, determine the similarity between the first table data and the second table data.
[0077] In implementation, the server can sum the similarities between the two sequences in the matched sequence pair and perform normalization processing on the sum value. The result of normalization is the similarity between the first table data and the second table data.
[0078] In actual application, the specific processing method for extracting the summary data corresponding to each sequence in the first table data based on the data included in each sequence in the determined first table data in step S108 can be various. The feature data can include sequence header features. Correspondingly, the following provides an optional processing method, such as Figure 4 As shown, it can specifically include the processing steps S1082 to S1086 below.
[0079] In step S1082, determine the first data feature according to the data in the target sequence in the determined first table data except for the first data.
[0080] Among them, the target sequence is any one sequence in the determined first table data.
[0081] In step S1084, determine the second data feature according to the first data in the target sequence in the determined first table data.
[0082] In step S1086, perform a consistency judgment on the first data feature and the second data feature, and in the case where the first data feature and the second data feature are consistent, determine the sequence header feature corresponding to the target sequence according to the first data in the target sequence.
[0083] In implementation, taking the preset arrangement direction as vertical arrangement as an example, for table data, there may be two situations for the column headers: one is that the table has column headers, and usually the first row of data is the column header; the other is that the table does not have column headers. Therefore, the key to extracting the sequence header features in the table data is to judge whether the first row of data in the table data is the column header.
[0084] Since the data types and data lengths of the data in the same column are consistent, the consistency characteristics (i.e., the first data feature) of the data in this column can be obtained by counting the data types and data lengths of the data from the second row to the last row of the target sequence. Then, the second data feature of the first row of data in this column is judged, and further whether the first data feature and the second data feature meet the consistency. If the consistency is met, there is no sequence header in this column; otherwise, it is considered that the first row of data is the sequence header, that is, the first data of the target sequence can be the sequence header feature of this sequence.
[0085] In practical applications, the specific processing method of extracting the summary of each sequence in the first table data based on the data included in each sequence determined in step S108 can be various. The feature data can include entity type features. Correspondingly, the following provides an optional processing method, such as Figure 5 shown, and it can specifically include the processing of the following step S1088.
[0086] In step S1088, obtain the preset named entities corresponding to the target service, and perform named entity recognition processing on each sequence in the first table data according to the preset named entities and the data included in each sequence in the first table data, so as to obtain the entity type features corresponding to each sequence in the first table data.
[0087] In implementation, the server can set different preset named entities for the target services in different scenarios. For example, for the target service in the financial scenario, the corresponding preset named entities can include entities such as resource transfer quantity, transaction type, and transaction subject; for the target service in the traffic scenario, the corresponding preset named entities can include entities such as name, user identifier, and train number identifier.
[0088] The server can construct corresponding entity recognizers for different preset named entities (such as recognizers constructed by natural language processing (NLP) algorithms). Furthermore, according to the entity recognizer of the preset named entity corresponding to the target service, and based on the data included in each sequence in the first table data, the matching scores corresponding to each preset named entity can be determined. Then, the server can determine the entity type features corresponding to each sequence according to the matching scores.
[0089] In practical applications, the specific processing method of extracting the summary of each sequence in the first table data based on the data included in each sequence determined in step S108 can be various. The feature data can include keyword features. Correspondingly, the following provides an optional processing method, such asFigure 6 As shown, it can specifically include the processing of the following steps S10810 to S10812.
[0090] In step S10810, word segmentation is performed on each data included in the target sequence in the first table data, and the occurrence frequency of each word obtained by the word segmentation is acquired.
[0091] Among them, the target sequence can be any sequence in the determined first table data.
[0092] In implementation, the server can use a pre-trained word segmentation model to perform word segmentation on each data included in the target sequence in the first table data, obtaining multiple words, and counting the occurrence frequency of each word. Among them, the word segmentation model can be a model constructed based on a preset deep learning algorithm.
[0093] In step S10812, according to the occurrence frequency of each word, screening processing is performed on the words to obtain the keyword features corresponding to the target sequence in the first table data.
[0094] In implementation, the server can perform screening processing on the words based on the preset frequency threshold according to the occurrence frequency of each word to obtain the keyword features. Or, the server can also perform sorting processing on the words according to the occurrence frequency of each word, and perform screening processing on the words according to the sorted words to obtain the keyword features. For example, the top ten words in the sorted words can be determined as the keyword features, etc.
[0095] In summary, taking the feature data including data type, data quantity, maximum value, minimum value, median, high-frequency value, sequence header, entity type, and keywords as an example, by compressing the sequence into a dictionary (i.e., the data of the aunt) that only contains 9 feature fields, the space complexity of table storage can be significantly reduced.
[0096] Suppose a table data with m rows and n columns, where m >> n. The space required to store this table data is mn. The storage space of the summary data generated by the sequence summary algorithm only requires 9n, and the compression ratio is m / 9. That is, the more the number of sequences in the table data, the greater the compression ratio.
[0097] Suppose the first table data is a table data with m rows and n columns, the second table data is a table data with M rows and N columns, and m < M, and the preset arrangement direction is vertical arrangement. The time complexity benefits brought by the sequence summary algorithm are analyzed for each stage of table similarity calculation below.
[0098] 1. Sequence similarity calculation. If the similarity between any two sequences in the first table data and the second table data is directly calculated, the time complexity of the similarity calculation for two sequences (with data volumes of m*1 and M*1 respectively) is O(m). While the time complexity of the similarity calculation for the corresponding digest data of the sequences (with data volumes of 9*1 and 9*1 respectively) is O(9). The efficiency of calculating similarity through the digest data of the sequences can be increased by m / 9 times.
[0099] 2. Table similarity calculation. After obtaining the similarity between any two sequences of the two table data, the Hungarian algorithm can be used to obtain the matching sequence pairs, and then, based on the similarity between the two sequences in the matching sequence pairs, the similarity between the first table data and the second table data is calculated. The time complexity of the Hungarian algorithm is O(n^3), and the time complexity benefit of the table digest algorithm here is 1.
[0100] By integrating the sequence similarity calculation and the table similarity calculation, it can be obtained that the final time complexity benefit of the digest algorithm is m / 9 * 1 = m / 9, that is, the digest algorithm can improve the efficiency of table similarity calculation, and the magnification of improvement can be m / 9. In addition, the more the number of sequences in the first table data and the second table data, the greater the magnification of efficiency improvement.
[0101] In practical applications, the specific processing methods for determining the first table data and the second table data corresponding to the table analysis request in step S104 can be various. The target service can be a risk detection service for a preset service. Correspondingly, the following provides an optional processing method, such as Figure 7 shown, which can specifically include the processing of the following steps S1042 to S1044.
[0102] In step S1042, according to the historical business data of the preset service, the first table data is determined.
[0103] In implementation, for example, taking the preset service as a resource transfer service, the historical business data of the preset service can be resource transfer business data without risks. Among them, the resource transfer business data can include multiple data such as resource transfer time, resource transfer quantity, resource transfer user, resource transfer platform, and user equipment.
[0104] In addition, in the case of a large amount of historical business data, the server can perform splitting processing on the historical business data to obtain multiple sub-business data, and construct the corresponding first table data according to each sub-business data, that is, the first table data determined by the server can be multiple.
[0105] For example, historical business data may include resource transfer data for multiple dates. For example, historical business data may include resource transfer data for the past six months. The server may split the historical business data on a quarterly basis to obtain two sub-business data, and respectively construct first table data corresponding to each sub-business data.
[0106] In step S1044, obtain the target business data corresponding to the currently executed preset business, and determine the second table data according to the target business data.
[0107] In implementation, taking the preset business as the resource transfer business as an example, the target business data may be the data required for executing the resource transfer business. For example, the target business data may include resource transfer time, resource transfer quantity, resource transfer user, resource transfer platform, user equipment, etc.
[0108] In this way, the server can determine whether there is a risk in executing the preset business through the table analysis results of the first table data and the second table data. That is, when the first table data does not contain risk data, the server can judge whether there is a risk in executing the preset business according to the similarity between the first table data and the second table data.
[0109] In practical applications, the specific processing methods for determining the first table data and the second table data corresponding to the table analysis request in step S104 can be various. The target business may be the product recommendation business. Correspondingly, the following provides an optional processing method, such as Figure 8 As shown, it may specifically include the processing of the following steps S1046~S1048.
[0110] In step S1046, determine the first table data according to the retrieval data input by the user.
[0111] In implementation, the server may perform keyword extraction processing on the retrieval data input by the user, and determine the first table data according to the extracted keywords. For example, the extracted keywords may include product name, product category, product price, product providing method, product timeliness, etc.
[0112] In step S1048, obtain the product data corresponding to the candidate products, and determine the second table data corresponding to each candidate product according to the product data corresponding to each candidate product.
[0113] Among them, the candidate products may be any recommendable products, and the product data corresponding to the candidate products may include data such as product name, product category, product price, product providing method, product timeliness, etc.
[0114] In implementation, to improve the efficiency of product recommendation, when the product category is included in the first table data, the server can screen candidate products according to the product category, and determine the second table data based on the product data corresponding to the screened candidate products.
[0115] In practical applications, the specific processing method for executing the target service according to the table analysis result in step S112 can be various. The target service can be a product recommendation service. Correspondingly, an optional processing method is provided below. For example, Figure 8 as shown, it can specifically include the processing of the following steps S1122 to S1124.
[0116] In step S1122, according to the similarity between the first table data and each second table data, the target table data is screened out from the second table data.
[0117] In implementation, the server can screen out the target table data from the second table data according to the preset product similarity threshold and the similarity between the first table data and each second table data. Or, the server can also sort the second table data according to the similarity between the first table data and each second table data, and determine the target table data based on the sorted second table data. For example, the first 5 in the sorted second table data can be determined as the target table data.
[0118] The screening method of the target table data can also be various, and different screening methods can be selected according to different actual application scenarios. This specification does not make specific limitations on this.
[0119] In step S1124, according to the product corresponding to the target table data, the product recommendation result corresponding to the product recommendation service is determined.
[0120] In implementation, the server can determine the product corresponding to the target table data as the product recommendation result corresponding to the product recommendation service.
[0121] In practical applications, the specific processing method for executing the target service according to the table analysis result in step S112 can be various. The target service can be a database cleaning service. The first table data and the second table data can be any two different table data in the preset database. Correspondingly, an optional processing method is provided below. For example, Figure 9 as shown, it can specifically include the processing of the following steps S1126 to S1128.
[0122] In step S1126, according to the similarity between the first table data and the second table data, the table data in the database is clustered to obtain multiple classes.
[0123] In step S1128, the data included in the tabular data corresponding to each class is de-duplicated to obtain a cleaned database.
[0124] In implementation, for each class, the server may summarize the data included in the tabular data corresponding to each class, perform duplicate checking on the summarized data. If duplicate sequences are found, the duplicate sequences may be de-duplicated to obtain a cleaned database.
[0125] The embodiments of this specification provide a data processing method. By receiving a tabular analysis request for a target service, in response to the tabular analysis request, determining first tabular data and second tabular data corresponding to the tabular analysis request, and adjusting the arrangement manner of the tabular content of the first tabular data and the second tabular data to a preset arrangement manner. According to the adjusted arrangement manner of the first tabular data and the second tabular data, determining the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement manner of the first tabular data and the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement manner of the second tabular data. Based on the data included in each sequence of the determined first tabular data, performing summary extraction processing on each sequence in the first tabular data to obtain summary data corresponding to each sequence in the first tabular data, and based on the data included in each sequence of the determined second tabular data, performing summary extraction processing on each sequence in the second tabular data to obtain summary data corresponding to each sequence in the second tabular data. According to the summary data corresponding to each sequence in the first tabular data and the summary data corresponding to each sequence in the second tabular data, determining the similarity between the first tabular data and the second tabular data. According to the similarity between the first tabular data and the second tabular data, determining the tabular analysis result, and according to the tabular analysis result, executing the target service. In this way, first, the arrangement manner of the tabular content of the tabular data to be analyzed (i.e., the first tabular data and the second tabular data) can be adjusted to a preset arrangement manner to improve the processing efficiency of subsequent tabular analysis. Second, for each tabular data to be analyzed, summary extraction processing is performed in units of sequences. When the amount of data included in the tabular data to be analyzed is large, while retaining the key information included in each sequence through the summary data corresponding to each sequence, the data volume can be compressed. Furthermore, through the summary data corresponding to each sequence, the similarity between the two tabular data to be analyzed can be quickly and accurately determined, reducing the time complexity of tabular analysis, improving the tabular analysis efficiency and accuracy, and improving the subsequent service processing efficiency.
[0126] The above is the data processing method provided by the embodiments of this specification. Based on the same idea, the embodiments of this specification also provide a data processing device, as Figure 10 shown.
[0127] The data processing device includes: a request receiving module 1001, a table determining module 1002, a data determining module 1003, an abstract extracting module 1004, a table matching module 1005, and a service execution module 1006, where: The request receiving module 1001 is configured to receive a table analysis request for a target service; The table determining module 1002 is configured to, in response to the table analysis request, determine first table data and second table data corresponding to the table analysis request, and adjust the arrangement manner of the table contents of the first table data and the second table data to a preset arrangement manner; The data determining module 1003 is configured to, according to the adjusted arrangement manner of the first table data and the second table data, determine the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement manner in the first table data and the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement manner in the second table data; The abstract extracting module 1004 is configured to perform an abstract extraction process on each sequence in the first table data based on the data included in each sequence determined in the first table data, to obtain abstract data corresponding to each sequence in the first table data, and perform an abstract extraction process on each sequence in the second table data based on the data included in each sequence determined in the second table data, to obtain abstract data corresponding to each sequence in the second table data; The table matching module 1005 is configured to determine the similarity between the first table data and the second table data according to the abstract data corresponding to each sequence in the first table data and the abstract data corresponding to each sequence in the second table data; The service execution module 1006 is configured to determine a table analysis result according to the similarity between the first table data and the second table data, and execute the target service according to the table analysis result.
[0128] In an embodiment of the present specification, the arrangement manner of the table contents of the first table data is determined according to the data type, data length, and data frequency of the data in the target row in the first table data, and the data type, data length, and data frequency of the data in the target column in the first table data, where the target row is any row in the first table data, and the target column is any column in the first table data.
[0129] In an embodiment of the present specification, the table matching module 1005 is configured to: Determine the sequence similarity between any two sequences in the first table data and the second table data according to the abstract data corresponding to each sequence in the first table data and the abstract data corresponding to each sequence in the second table data; Perform sequence matching processing on the first table data and the second table data according to the sequence similarity between any two sequences in the first table data and the second table data, and determine the matching sequence pairs in the first table data and the second table data; Determine the similarity between the first table data and the second table data according to the similarity between the two sequences in the matching sequence pairs.
[0130] In the embodiments of this specification, the abstract data includes multiple feature data, and the table matching module 1005 is configured to: Determine the similarity between any two sequences in the first table data and the second table data for each feature data according to each feature data in the abstract data corresponding to each sequence in the first table data and each feature data in the abstract data corresponding to each sequence in the second table data; Determine the sequence similarity between any two sequences in the first table data and the second table data according to the preset weight value corresponding to each feature data and the similarity between any two sequences in the first table data and the second table data for each feature data.
[0131] In the embodiments of this specification, the device further includes: A data acquisition module, configured to acquire sample sequence data corresponding to the target service, first sequence data having a preset similarity relationship with the sample sequence data, and second sequence data having no preset similarity relationship with the sample sequence data; A first determination module, configured to respectively determine a first similarity between the sample sequence data and each first sequence data, and a second similarity between the sample feature data and each second sample data according to the abstract data corresponding to the sample sequence data, the abstract data corresponding to the first sequence data, the abstract data corresponding to the second sequence data, and the initial weight value corresponding to each feature data; A second determination module, configured to determine whether the initial weight value meets a preset screening requirement according to the first similarity and the second similarity, and determine the initial weight value as the preset weight value corresponding to each feature data when the initial weight value meets the preset screening requirement; A third determination module, configured to, when the initial weight value does not meet the preset screening requirements, adjust the initial weight value, and continue to determine the first similarity and the second similarity according to the adjusted initial weight value until the adjusted initial weight value meets the preset screening requirements, and determine the adjusted initial weight value as the preset weight value corresponding to each piece of the feature data.
[0132] In the embodiments of the present specification, the feature data includes sequence header features, and the abstract extraction module 1004 is configured to: Determine first data features according to the data other than the first data in the target sequence in the determined first table data, where the target sequence is any one of the sequences in the determined first table data; Determine second data features according to the first data in the target sequence in the determined first table data; Perform a consistency judgment on the first data features and the second data features, and when there is consistency between the first data features and the second data features, determine the sequence header features corresponding to the target sequence according to the first data in the target sequence.
[0133] In the embodiments of the present specification, the feature data includes entity type features, and the abstract extraction module 1004 is configured to: Obtain a preset named entity corresponding to the target service, and perform named entity recognition processing on each sequence in the first table data according to the preset named entity and the data included in each sequence in the first table data, to obtain the entity type features corresponding to each sequence in the first table data.
[0134] The feature data includes keyword features, and the abstract extraction module is configured to: Perform word segmentation processing on each data included in the target sequence in the first table data, and obtain the occurrence frequency of each word obtained by the word segmentation processing, where the target sequence is any one of the sequences in the determined first table data; Perform screening processing on the words according to the occurrence frequency of each word, to obtain the keyword features corresponding to the target sequence in the first table data.
[0135] In the embodiments of the present specification, the target service is a risk detection service for a preset service, and the table determination module 1002 is configured to: Determine the first table data according to the historical service data of the preset service; Obtain target service data corresponding to the currently executed preset service, and determine the second table data according to the target service data.
[0136] In the embodiments of this specification, the target service is a product recommendation service. The table determination module 1002 is configured to: Determine the first table data according to the retrieval data input by the user; Obtain the product data corresponding to the candidate products, and determine the second table data corresponding to each candidate product according to the product data corresponding to each candidate product; The service execution module 1006 is configured to: Screen out the target table data from the second table data according to the similarity between the first table data and each second table data; Determine the product recommendation result corresponding to the product recommendation service according to the product corresponding to the target table data.
[0137] In the embodiments of this specification, the target service is a database cleaning service. The first table data and the second table data are any two different table data in a preset database. The service execution module 1006 is configured to: Perform clustering processing on the table data of the database according to the similarity between the first table data and the second table data to obtain multiple classes; Perform duplicate removal processing on the data included in the table data corresponding to each class to obtain a cleaned database.
[0138] An embodiment of this specification provides a data processing device. By receiving a form analysis request for a target service, in response to the form analysis request, determining first form data and second form data corresponding to the form analysis request, and adjusting the arrangement manner of the form content of the first form data and the second form data to a preset arrangement manner. According to the adjusted arrangement manner of the first form data and the second form data, determining the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the first form data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the second form data. Based on the data included in each sequence determined in the first form data, performing a summary extraction process on each sequence in the first form data to obtain summary data corresponding to each sequence in the first form data, and based on the data included in each sequence determined in the second form data, performing a summary extraction process on each sequence in the second form data to obtain summary data corresponding to each sequence in the second form data. According to the summary data corresponding to each sequence in the first form data and the summary data corresponding to each sequence in the second form data, determining the similarity between the first form data and the second form data. According to the similarity between the first form data and the second form data, determining a form analysis result, and according to the form analysis result, executing the target service. In this way, first, the arrangement manner of the form content of the form data to be analyzed (i.e., the first form data and the second form data) can be adjusted to a preset arrangement manner to improve the processing efficiency of subsequent form analysis. Second, for each form data to be analyzed, a summary extraction process is performed in units of sequences. When the amount of data included in the form data to be analyzed is large, while retaining the key information included in each sequence through the summary data corresponding to each sequence, the data amount can be compressed. Furthermore, through the summary data corresponding to each sequence, the similarity between the two form data to be analyzed can be quickly and accurately determined, reducing the time complexity of form analysis, improving the form analysis efficiency and accuracy, and improving the subsequent service processing efficiency.
[0139] The above is the data processing device provided by the embodiment of this specification. Based on the same idea, the embodiment of this specification also provides a data processing device, as Figure 11 shown.
[0140] The data processing device may be a terminal device or a server provided in the above embodiment, etc.
[0141] Data processing devices can vary significantly in configuration or performance and may include one or more processors 1101 and a memory 1102. One or more application programs or data may be stored in the memory 1102. Among them, the memory 1102 can be transient storage or persistent storage. The application programs stored in the memory 1102 may include one or more modules (not shown in the figure), and each module may include a series of computer-executable instructions in the data processing device. Further, the processor 1101 can be set to communicate with the memory 1102 and execute a series of computer-executable instructions in the memory 1102 on the data processing device. The data processing device may also include one or more power supplies 1103, one or more wired or wireless network interfaces 1104, one or more input / output interfaces 1105, and one or more keyboards 1106.
[0142] Specifically, in this embodiment, the data processing device includes a memory and one or more programs. One or more of the programs are stored in the memory, and one or more of the programs may include one or more modules. Each module may include a series of computer-executable instructions in the data processing device and is configured to be executed by one or more processors. The one or more programs include the following computer-executable instructions: Receive a form analysis request for a target service; In response to the form analysis request, determine first form data and second form data corresponding to the form analysis request, and adjust the arrangement of the form content of the first form data and the second form data to a preset arrangement; According to the adjusted arrangement of the first form data and the second form data, determine the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement of the first form data and the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement of the second form data; Based on the data included in each sequence of the determined first form data, perform a digest extraction process on each sequence of the first form data to obtain the digest data corresponding to each sequence of the first form data, and based on the data included in each sequence of the determined second form data, perform a digest extraction process on each sequence of the second form data to obtain the digest data corresponding to each sequence of the second form data; According to the digest data corresponding to each sequence of the first form data and the digest data corresponding to each sequence of the second form data, determine the similarity between the first form data and the second form data; Determine a table analysis result according to the similarity between the first table data and the second table data, and execute the target service according to the table analysis result.
[0143] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between each embodiment, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiment of the data processing device, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.
[0144] The embodiment of this specification provides a data processing device. By receiving a table analysis request for a target service, in response to the table analysis request, determine the first table data and the second table data corresponding to the table analysis request, and adjust the arrangement manner of the table content of the first table data and the second table data to a preset arrangement manner. According to the adjusted arrangement manner of the first table data and the second table data, determine the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner of the first table data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner of the second table data. Based on the data included in each sequence of the determined first table data, perform a summary extraction process on each sequence in the first table data to obtain the summary data corresponding to each sequence in the first table data, and based on the data included in each sequence of the determined second table data, perform a summary extraction process on each sequence in the second table data to obtain the summary data corresponding to each sequence in the second table data. According to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data, determine the similarity between the first table data and the second table data. According to the similarity between the first table data and the second table data, determine the table analysis result, and execute the target service according to the table analysis result. In this way, first, the arrangement manner of the table content of the table data to be analyzed (i.e., the first table data and the second table data) can be adjusted to a preset arrangement manner to improve the processing efficiency of subsequent table analysis. Second, for each table data to be analyzed, perform a summary extraction process in units of sequences. When the amount of data included in the table data to be analyzed is large, while retaining the key information included in each sequence through the summary data corresponding to each sequence, the data volume can be compressed. Furthermore, through the summary data corresponding to each sequence, the similarity between the two table data to be analyzed can be quickly and accurately determined, reducing the time complexity of table analysis, improving the efficiency and accuracy of table analysis, and improving the efficiency of subsequent service processing.
[0145] Further, based on the above Figures 1 to 9For the method shown, one or more embodiments of this specification also provide a storage medium for storing computer-executable instruction information. In a specific embodiment, the storage medium can be a USB flash drive, an optical disc, a hard disk, etc. When the computer-executable instruction information stored in the storage medium is executed by a processor, the following processes can be implemented: Receive a table analysis request for a target service; In response to the table analysis request, determine first table data and second table data corresponding to the table analysis request, and adjust the arrangement manner of the table content of the first table data and the second table data to a preset arrangement manner; According to the adjusted arrangement manner of the first table data and the second table data, determine the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the first table data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the second table data; Based on the data included in each sequence in the determined first table data, perform a summary extraction process on each sequence in the first table data to obtain summary data corresponding to each sequence in the first table data, and based on the data included in each sequence in the determined second table data, perform a summary extraction process on each sequence in the second table data to obtain summary data corresponding to each sequence in the second table data; According to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data, determine the similarity between the first table data and the second table data; According to the similarity between the first table data and the second table data, determine a table analysis result, and execute the target service according to the table analysis result.
[0146] The various embodiments in this specification are all described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the above storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0147] An embodiment of this specification provides a storage medium. By receiving a table analysis request for a target service, in response to the table analysis request, determining first table data and second table data corresponding to the table analysis request, and adjusting the arrangement mode of the table content of the first table data and the second table data to a preset arrangement mode. According to the adjusted arrangement modes of the first table data and the second table data, determining the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement mode in the first table data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement mode in the second table data. Based on the data included in each sequence determined in the first table data, performing a summary extraction process on each sequence in the first table data to obtain summary data corresponding to each sequence in the first table data, and based on the data included in each sequence determined in the second table data, performing a summary extraction process on each sequence in the second table data to obtain summary data corresponding to each sequence in the second table data. According to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data, determining the similarity between the first table data and the second table data. According to the similarity between the first table data and the second table data, determining the table analysis result, and according to the table analysis result, executing the target service. In this way, first, the arrangement mode of the table content of the table data to be analyzed (i.e., the first table data and the second table data) can be adjusted to a preset arrangement mode to improve the processing efficiency of subsequent table analysis. Second, for each table data to be analyzed, a summary extraction process is performed in units of sequences. When the amount of data included in the table data to be analyzed is large, while retaining the key information included in each sequence through the summary data corresponding to each sequence, the data volume can be compressed. Furthermore, through the summary data corresponding to each sequence, the similarity between the two table data to be analyzed can be quickly and accurately determined, reducing the time complexity of table analysis, improving the efficiency and accuracy of table analysis, and improving the efficiency of subsequent service processing.
[0148] Further, based on the above Figures 1 to 9 method shown, one or more embodiments of this specification also provide a computer program product, including a computer program. When the computer program in this computer program product is executed by a processor, the following processes can be implemented: Receiving a table analysis request for a target service; In response to the table analysis request, determining first table data and second table data corresponding to the table analysis request, and adjusting the arrangement mode of the table content of the first table data and the second table data to a preset arrangement mode; Determine the data included in each sequence in the arranged direction corresponding to the adjusted arrangement mode in the first tabular data and the data included in each sequence in the arranged direction corresponding to the adjusted arrangement mode in the second tabular data according to the adjusted arrangement mode based on the first tabular data and the second tabular data; Based on the data included in each sequence in the determined first tabular data, perform abstract extraction processing on each sequence in the first tabular data to obtain the abstract data corresponding to each sequence in the first tabular data, and based on the data included in each sequence in the determined second tabular data, perform abstract extraction processing on each sequence in the second tabular data to obtain the abstract data corresponding to each sequence in the second tabular data; Determine the similarity between the first tabular data and the second tabular data according to the abstract data corresponding to each sequence in the first tabular data and the abstract data corresponding to each sequence in the second tabular data; Determine the table analysis result according to the similarity between the first tabular data and the second tabular data, and execute the target service according to the table analysis result.
[0149] Each embodiment in this specification is described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the above embodiment of a computer program product, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0150] An embodiment of this specification provides a computer program product. By receiving a form analysis request for a target service, in response to the form analysis request, first table data and second table data corresponding to the form analysis request are determined, and the arrangement of the form content of the first table data and the second table data is adjusted to a preset arrangement. According to the adjusted arrangement of the first table data and the second table data, the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement in the first table data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement in the second table data are determined. Based on the data included in each sequence determined in the first table data, a summary extraction process is performed on each sequence in the first table data to obtain summary data corresponding to each sequence in the first table data, and based on the data included in each sequence determined in the second table data, a summary extraction process is performed on each sequence in the second table data to obtain summary data corresponding to each sequence in the second table data. According to the summary data corresponding to each sequence in the first table data and the summary data corresponding to each sequence in the second table data, the similarity between the first table data and the second table data is determined. According to the similarity between the first table data and the second table data, a form analysis result is determined, and according to the form analysis result, the target service is executed. In this way, first, the arrangement of the form content of the form data to be analyzed (i.e., the first table data and the second table data) can be adjusted to a preset arrangement to improve the processing efficiency of subsequent form analysis. Second, for each form data to be analyzed, a summary extraction process is performed in units of sequences. When the amount of data included in the form data to be analyzed is large, while retaining the key information included in each sequence through the summary data corresponding to each sequence, the data volume can be compressed. Furthermore, through the summary data corresponding to each sequence, the similarity between the two form data to be analyzed can be quickly and accurately determined, reducing the time complexity of form analysis, improving the form analysis efficiency and accuracy, and improving the subsequent service processing efficiency.
[0151] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0152] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0153] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0154] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0155] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing one or more embodiments of this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0156] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0157] Embodiments of the present specification are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present specification. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable parallel and serial devices for fraud cases to generate a machine, such that the instructions executed by the processor of the computer or other programmable parallel and serial devices for fraud cases generate means for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0158] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable parallel and serial devices for fraud cases to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0159] These computer program instructions can also be loaded onto a computer or other programmable parallel and serial devices for fraud cases, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in the process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0160] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0161] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0162] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined in this article, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0163] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0164] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, one or more embodiments of this specification may be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Furthermore, one or more embodiments of this specification may be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0165] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0166] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment.
[0167] The above is only the embodiment of this specification and is not used to limit this document. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A data processing method, comprising: Receiving a table analysis request for a target service; In response to the table analysis request, determining first table data and second table data corresponding to the table analysis request, and adjusting the arrangement manner of the table content of the first table data and the second table data to a preset arrangement manner; According to the adjusted arrangement manner of the first table data and the second table data, determining the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the first table data and the data included in each sequence in the arrangement direction corresponding to the adjusted arrangement manner in the second table data; Based on the data included in each sequence in the determined first table data, performing a digest extraction process on each sequence in the first table data to obtain the digest data corresponding to each sequence in the first table data, and based on the data included in each sequence in the determined second table data, performing a digest extraction process on each sequence in the second table data to obtain the digest data corresponding to each sequence in the second table data; According to the digest data corresponding to each sequence in the first table data and the digest data corresponding to each sequence in the second table data, determining the similarity between the first table data and the second table data; According to the similarity between the first table data and the second table data, determining a table analysis result, and according to the table analysis result, executing the target service.
2. The method according to claim 1, wherein the arrangement of the table content of the first table data is determined according to the data type, data length, and data frequency of the data in the target row of the first table data, and the data type, data length, and data frequency of the data in the target column of the first table data, wherein, The target row is any row in the first table data, and the target column is any column in the first table data.
3. The method according to claim 2, wherein the determining the similarity between the first table data and the second table data according to the digest data corresponding to each sequence in the first table data and the digest data corresponding to each sequence in the second table data comprises: Determining the sequence similarity between any two sequences in the first table data and the second table data according to the digest data corresponding to each sequence in the first table data and the digest data corresponding to each sequence in the second table data; Performing a sequence matching process on the first table data and the second table data according to the sequence similarity between any two sequences in the first table data and the second table data, and determining the matching sequence pairs in the first table data and the second table data; Determining the similarity between the first table data and the second table data according to the similarity between the two sequences in the matching sequence pairs.
4. The method according to claim 3, wherein the digest data includes multiple feature data, and the determining the sequence similarity between any two sequences in the first table data and the second table data according to the digest data corresponding to each sequence in the first table data and the digest data corresponding to each sequence in the second table data comprises: Determine the similarity between any two sequences in the first table data and the second table data for each feature data according to each feature data in the summary data corresponding to each sequence in the first table data and each feature data in the summary data corresponding to each sequence in the second table data; Determine the sequence similarity between any two sequences in the first table data and the second table data according to the preset weight value corresponding to each feature data and the similarity between any two sequences in the first table data and the second table data for each feature data; 5. The method according to claim 4, before determining the sequence similarity between any two sequences in the first table data and the second table data according to the preset weight value corresponding to each feature data and the similarity between any two sequences in the first table data and the second table data for each feature data, further includes: Obtain sample sequence data corresponding to the target service, first sequence data having a preset similarity relationship with the sample sequence data, and second sequence data having no preset similarity relationship with the sample sequence data; According to the summary data corresponding to the sample sequence data, the summary data corresponding to the first sequence data, the summary data corresponding to the second sequence data, and the initial weight value corresponding to each feature data, respectively determine the first similarity between the sample sequence data and each first sequence data, and the second similarity between the sample feature data and each second sample data; According to the first similarity and the second similarity, determine whether the initial weight value meets the preset screening requirements. When the initial weight value meets the preset screening requirements, determine the initial weight value as the preset weight value corresponding to each feature data; When the initial weight value does not meet the preset screening requirements, perform adjustment processing on the initial weight value, and continue to determine the first similarity and the second similarity according to the adjusted initial weight value until the adjusted initial weight value meets the preset screening requirements, and determine the adjusted initial weight value as the preset weight value corresponding to each feature data.
6. The method according to claim 4, wherein the feature data includes a sequence header feature, and the extracting summary data corresponding to each sequence in the first table data by performing summary extraction processing on each sequence in the first table data based on the data included in each sequence in the determined first table data includes: Determine a first data feature according to the data other than the first data in the target sequence in the determined first table data, where the target sequence is any one sequence in the determined first table data; Determine a second data feature according to the first data in the target sequence in the determined first table data; Perform a consistency judgment on the first data feature and the second data feature, and when there is consistency between the first data feature and the second data feature, determine the sequence header feature corresponding to the target sequence according to the first data in the target sequence.
7. The method according to claim 4, wherein the feature data includes an entity type feature, and the step of performing a summary extraction process on each sequence in the first table data based on the data included in each sequence in the determined first table data to obtain the summary data corresponding to each sequence in the first table data includes: Obtain a preset named entity corresponding to the target service, and perform a named entity recognition process on each sequence in the first table data according to the preset named entity and the data included in each sequence in the first table data, so as to obtain the entity type feature corresponding to each sequence in the first table data.
8. The method according to claim 4, wherein the feature data includes a keyword feature, and the step of performing a summary extraction process on each sequence in the first table data based on the data included in each sequence in the determined first table data to obtain the summary data corresponding to each sequence in the first table data includes: Perform word segmentation on each data included in the target sequence in the first table data, and obtain the occurrence frequency of each word obtained by the word segmentation, where the target sequence is any sequence in the determined first table data; Perform a screening process on the words according to the occurrence frequency of each word to obtain the keyword feature corresponding to the target sequence in the first table data.
9. The method according to claim 1, wherein the target service is a risk detection service for a preset service, and the step of determining the first table data and the second table data corresponding to the table analysis request includes: Determine the first table data according to the historical service data of the preset service; Obtain the target service data corresponding to the currently executed preset service, and determine the second table data according to the target service data.
10. The method according to claim 1, wherein the target service is a product recommendation service, and the step of determining the first table data and the second table data corresponding to the table analysis request includes: Determine the first table data according to the retrieval data input by the user; Obtain the product data corresponding to the candidate products, and determine the second table data corresponding to each candidate product according to the product data corresponding to each candidate product; The step of performing the target service according to the table analysis result includes: Screen out the target table data from the second table data according to the similarity between the first table data and each second table data; Determine the product recommendation result corresponding to the product recommendation service according to the product corresponding to the target table data.
11. According to the method described in claim 1, where the target service is a database cleaning service, and the first table data and the second table data are any two different table data in a preset database, the execution of the target service according to the table analysis result includes: Performing clustering processing on the table data of the database according to the similarity between the first table data and the second table data to obtain multiple classes; Performing deduplication processing on the data included in the table data corresponding to each class to obtain a cleaned database.
12. A data processing device, comprising: A request receiving module, configured to receive a table analysis request for a target service; A table determining module, configured to, in response to the table analysis request, determine first table data and second table data corresponding to the table analysis request, and adjust the arrangement manner of the table content of the first table data and the second table data to a preset arrangement manner; A data determining module, configured to determine the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement manner of the first table data and the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement manner of the second table data according to the adjusted arrangement manners of the first table data and the second table data; An abstract extraction module, configured to perform abstract extraction processing on each sequence in the first table data based on the data included in each sequence determined in the first table data to obtain abstract data corresponding to each sequence in the first table data, and perform abstract extraction processing on each sequence in the second table data based on the data included in each sequence determined in the second table data to obtain abstract data corresponding to each sequence in the second table data; A table matching module, configured to determine the similarity between the first table data and the second table data according to the abstract data corresponding to each sequence in the first table data and the abstract data corresponding to each sequence in the second table data; A service execution module, configured to determine a table analysis result according to the similarity between the first table data and the second table data, and execute the target service according to the table analysis result.
13. A data processing device, the data processing device includes: A processor; And A memory arranged to store computer-executable instructions, the executable instructions, when executed, cause the processor to: Receive a table analysis request for a target service; In response to the table analysis request, determine first table data and second table data corresponding to the table analysis request, and adjust the arrangement manner of the table content of the first table data and the second table data to a preset arrangement manner; Determine the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement manner of the first table data and the data included in each sequence in the corresponding arrangement direction of the adjusted arrangement manner of the second table data according to the adjusted arrangement manners of the first table data and the second table data; Based on the data included in each sequence in the determined first tabular data, perform abstract extraction processing on each sequence in the first tabular data to obtain the abstract data corresponding to each sequence in the first tabular data, and based on the data included in each sequence in the determined second tabular data, perform abstract extraction processing on each sequence in the second tabular data to obtain the abstract data corresponding to each sequence in the second tabular data; Determine the similarity between the first tabular data and the second tabular data according to the abstract data corresponding to each sequence in the first tabular data and the abstract data corresponding to each sequence in the second tabular data; Determine the table analysis result according to the similarity between the first tabular data and the second tabular data, and execute the target service according to the table analysis result.