Data processing method and device, medium and product

Through feature nearest neighbor filters, the business data of different data sources of the bank is integrated, which solves the problems of large data storage space requirements and cumbersome usage processes, realizes the integrity, reliability and scalability of data integration, and discovers hidden relationships between data.

CN120470293APending Publication Date: 2025-08-12AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510658367.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Bank data storage under different sources leads to large demand for data storage space and cumbersome usage processes.

Method used

A characteristic nearest neighbor filter is used to integrate the business data combination from different data sources to generate a target data set.

Benefits of technology

On the premise of ensuring data integrity, reliability and scalability, data from different data sources are integrated into a target data collection, and hidden relationships between data are discovered and analyzed and processed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470293A_ABST
    Figure CN120470293A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, a medium and a product, and belongs to the technical field of data processing. The method comprises the following steps: acquiring a first business data combination from a first data source and a second business data combination from a second data source in a preset time period; and integrating the first service data combination and the second service data combination based on a feature nearest neighbor filter to obtain a target data set. According to the technical scheme provided by the embodiment of the invention, the first business data combination and the second business data combination from different data sources can be integrated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a data processing method, device, medium and product. Background Art

[0002] Currently, banks usually store business data from different sources in different locations, which leads to a large demand for data storage space and a more cumbersome subsequent data usage process. Summary of the Invention

[0003] The present invention provides a data processing method, device, medium and product to solve the problem of cumbersome data use caused by the existing data storage mode.

[0004] According to one aspect of the present invention, there is provided a data processing method, comprising:

[0005] Acquire a first business data combination from a first data source and a second business data combination from a second data source within a predetermined time period;

[0006] The first service data combination and the second service data combination are integrated based on a feature nearest neighbor filter to obtain a target data set.

[0007] According to another aspect of the present invention, there is provided a data processing apparatus, comprising:

[0008] A data acquisition module, configured to acquire a first business data combination from a first data source and a second business data combination from a second data source within a predetermined time period;

[0009] An integration module is used to integrate the first business data combination and the second business data combination based on a feature nearest neighbor filter to obtain a target data set.

[0010] According to another aspect of the present invention, an electronic device is provided, comprising:

[0011] at least one processor; and

[0012] a memory communicatively connected to the at least one processor; wherein,

[0013] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the data processing method described in any embodiment of the present invention.

[0014] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data processing method according to any embodiment of the present invention when executed.

[0015] According to another aspect of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the data processing method according to any one of the embodiments is implemented.

[0016] The technical solution of the data processing method provided by the embodiment of the present invention is that since the first business data combination comes from the first data source and the second business data comes from the second data source, there may be differences in the fields between the first business data combination and the second business data combination. The first business data combination and the second business data combination are integrated from multiple dimensions based on the feature nearest neighbor filter, and the first business data combination and the second business data combination can be integrated into a target data set; the technical effect of integrating the first business data combination and the second business data combination into the target data set under the premise of ensuring data integrity, reliability and scalability is achieved, which is conducive to discovering hidden relationships between data and further analysis and processing.

[0017] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 is a flow chart of a data processing method provided according to an embodiment of the present invention;

[0020] Figure 2 is another flow chart of a data processing method provided according to an embodiment of the present invention;

[0021] Figure 3 is another flow chart of a data processing method provided according to an embodiment of the present invention;

[0022] Figure 4 is a structural diagram of a data processing device provided according to an embodiment of the present invention;

[0023] Figure 5is another structural diagram of a data processing device provided according to an embodiment of the present invention;

[0024] Figure 6 is another structural diagram of a data processing device provided according to an embodiment of the present invention;

[0025] Figure 7 It is a structural diagram of an electronic device for implementing the data processing method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] It should be noted that the collection, collection, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solutions disclosed herein all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken with respect to user personal information to prevent unauthorized access to user personal information data and to safeguard the security of user personal information, network security, and national security.

[0029] Figure 1 A flowchart of a data processing method is provided for an embodiment of the present invention. This embodiment is applicable to the case of automatically completing the integrated storage of business data from different data sources. The method can be executed by a data processing device, which can be implemented in the form of hardware and / or software. The data processing device can be configured in a processor of an electronic device. Figure 1 As shown, the method includes:

[0030] S110: Acquire a first business data combination from a first data source and a second business data combination from a second data source within a predetermined time period.

[0031] The scheduled time period is configured as a custom item, and users can set it according to actual needs. It can be configured to six months, one month, etc.

[0032] The first data source corresponds to the first type of business, and the second data source corresponds to the second type of business. For example, in a bank custody business scenario, the same user may conduct two types of business at the same bank, such as the first type of business and the second type of business. The first business data generated by the first type of business and the second business data generated by the second type of business will typically share some common fields, such as user information.

[0033] The first service data combination includes multiple first service data, and the second service data combination includes multiple second service data. Taking the first service data combination as an example, it includes first service data for multiple users, and for each of the multiple users, the first service data combination may include only one first service data or multiple first service data.

[0034] In one embodiment, in response to a data integration request from a terminal, required data is queried from a database in a mapping form, and preliminary processing is performed on the data, such as data merging, to generate data classification results and obtain a first business data combination and a second business data combination.

[0035] In one embodiment, some fields of each first business data in the first business data combination and each second business data in the second business data combination are the same. There is a high integration requirement between the business data with the same fields.

[0036] S120 . Integrate the first service data combination and the second service data combination based on a feature nearest neighbor filter to obtain a first data set.

[0037] Among them, Feature-Nearest Neighbor Filter (F-NNF) is a nonlinear filtering method based on nearest neighbor search (NNS).

[0038] Regarding the nearest neighbor search, given a point set P = {p1, p2, ..., pn} in an n-dimensional space and a query point q, the goal of the nearest neighbor search is to search for the point pi that is closest to the query point q from the point set P, where the distance can be Euclidean distance, Manhattan distance, cosine similarity or other similarity.

[0039] Specifically:

[0040] Step a1: for each first business data in the first business data combination, if the second business data combination contains second business data that meets the similarity condition with the current first business data, then add the second business data to the target data set.

[0041] In one embodiment, if the similarity offset between the current first business data and the second business data under the same identifier in the second business data combination is less than a predetermined offset threshold, it is considered that there is second business data in the second business data combination that meets the similarity conditions with the first business data, that is, the two have a high similarity. At this time, the second business data is added to the target data set to update the target data set.

[0042] Step a2: If the second business data combination does not contain any second business data that meets the similarity condition with the current first business data, determine the sum of the similarity offsets of other first business data in the first business data combination relative to the corresponding second business data in the second business data combination, and determine the local closest distance between the second business data combination and the current first business data based on the nearest neighbor search algorithm.

[0043] Based on the data neighbor relationship, the maximum similarity between the current first business data and the second business data with the same data identifier is determined. If the upper limit of the similarity is still below a predetermined threshold, the two are not considered for integrated analysis. In this case, the distance between the current first business data and the closest second business data in the second business data combination needs to be calculated.

[0044] Optionally, determine the total similarity offset between each first business data in the first business data combination and the corresponding second business data in the second business data combination, as well as the current similarity offset of the current first business data relative to the second business data with the same identifier in the second business data; and take the difference between the total similarity offset and the current similarity offset as the sum of the similarity offsets of other first business data in the first business data combination relative to the corresponding second business data in the second business data combination.

[0045] The total similarity offset is determined by the following formula:

[0046]

[0047] Among them, flag is the total similarity offset, i is the business data identifier, r i is the first business data identified as i in the first business data, k i It is the second business data identified as i in the second business data combination. is the similarity offset of the first business data identified as i relative to the second business data identified as i.

[0048] If the identifier of the current first business data is i, then the current similarity offset of the current first business data relative to the second business data with the same identifier in the second business data is

[0049] The sum of the above similarity offsets can be expressed as:

[0050] Among them, the parameterized description of the nearest neighbor search algorithm can be expressed as NNS(r,d,I), where r is the search radius or search range, that is, looking for neighboring points within a distance r; d is the dimension, which represents the dimension of the data point. The higher the dimension, the greater the computational complexity of the nearest neighbor search; I is the input second business data combination and the current first business data; its output result is the local nearest distance, which is specifically the distance between the target second business data and the current first business data, where the target second business data is the second business data that is closest to the current first business data in the second business data.

[0051] Step a3: The sum of the similarity offsets and the sum of the local closest distances are used as the current intermediate variable.

[0052] For example,

[0053] Step a4: If the current first business data belongs to the reference set but does not belong to the first business data combination, the sum of the current intermediate variable, the similarity offset and the local closest distance is used as the updated current intermediate variable.

[0054] Specifically, if the current first service data belongs to the reference set but does not belong to the first service data combination, the following steps are performed:

[0055]

[0056] The latest Temp is the updated current intermediate variable, that is, the maximum similarity that can be achieved between the current first business data and the second business data under the same identifier.

[0057] Step a5: When the updated current intermediate variable meets the boundary conditions, the sum of the current target data set and the updated current intermediate variable is used as the updated current target data set.

[0058] Among them, the boundary condition is a predetermined threshold. If the updated current intermediate variable is less than the predetermined threshold, the sum of the current target data set and the updated current intermediate variable is used as the updated current target data set; if the updated current intermediate variable is less than the predetermined threshold, the current intermediate variable is ignored.

[0059] The technical solution of the data processing method provided by the embodiment of the present invention is that since the first business data combination comes from the first data source and the second business data comes from the second data source, there may be differences in the fields between the first business data combination and the second business data combination. The first business data combination and the second business data combination are integrated from multiple dimensions based on the feature nearest neighbor filter, and the first business data combination and the second business data combination can be integrated into a target data set; the technical effect of integrating the first business data combination and the second business data combination into the target data set under the premise of ensuring data integrity, reliability and scalability is achieved, which is conducive to discovering hidden relationships between data and further analysis and processing.

[0060] Figure 2 This is another flow chart of the data processing method provided by an embodiment of the present invention. This embodiment adds a data export step based on the above embodiment. Figure 2 As shown, the method includes:

[0061] S210: Acquire a first business data combination from a first data source and a second business data combination from a second data source within a predetermined time period.

[0062] S220 . Integrate the first business data combination and the second business data combination based on a feature nearest neighbor filter to obtain a target data set.

[0063] S230: Determine at least one query subject in response to the data export request.

[0064] In one embodiment, a user clicks or touches a data export option on an export interface of a terminal. The middleware server, in response to the user's data export operation on the export page, determines a data export request and at least one query subject corresponding to the data export request. The data export request includes information such as a request header and a request body.

[0065] Exemplarily, the data export request includes an export identifier, and at least one query subject corresponding to the current data export request is determined based on a predetermined correspondence between the export identifier and a query subject combination.

[0066] S240: Create a target table, where the target table includes at least one sub-table, and a subject name of the at least one sub-table corresponds to the at least one query subject in a one-to-one manner.

[0067] After the at least one query subject is determined, a target table is created. The target table includes subtables corresponding to the query subjects. The subject name of each subtable can be the corresponding query subject, or includes the corresponding query subject.

[0068] Exemplarily, the at least one query subject includes T1 and T2. The target table includes two sub-tables, one sub-table has a subject name of T1 and the other sub-table has a subject name of T2; or,

[0069] The subject name of one subtable is T1 table, and the subject name of the other subtable is T2 table.

[0070] In one embodiment, a workbook object is created based on the ExcelUtils tool class, and then worksheets corresponding to the query subjects in the at least one query subject are created in the workbook object. The initSheet() method is then called to initialize each worksheet and create a subject name in each worksheet so that the subject name of each worksheet includes the corresponding query subject.

[0071] S250: Export the data in the target data set corresponding to each query subject in the at least one query subject to a subtable corresponding to the subject name in the target table to update the target table.

[0072] Specifically, based on pre-created query subject and field combinations, the corresponding field combinations for each query subject are determined. Based on the corresponding field combinations for each query subject, the corresponding data is pulled from the target data set and stored in the subtable corresponding to the query subject, thereby updating the target table. The updated target table is then delivered to the terminal page in the form of a response body.

[0073] In one embodiment, after the target table is updated, a table configuration interface is displayed; the table name of the target table is determined according to the name configuration information received by the table configuration interface, and the storage path of the target table is determined according to the storage configuration information received by the table configuration interface; the target table is named as the table name, and the storage of the target table is completed according to the storage path.

[0074] Specifically, the table configuration interface includes a name configuration box and a storage path configuration box. The table name of the target table is determined according to the name configuration information entered by the user in the name configuration box, and the storage path of the target table is determined according to the address configuration information entered by the user in the address configuration box or the target address selected; then the target table is named as the table name, and the target table is stored under the storage path.

[0075] In one embodiment, the response.getOutputStream() and workbook.write() methods are used to export the file through an IO stream, and the download result for the target table is displayed on the terminal page.

[0076] The technical solution provided by the embodiment of the present invention exports data of various dimensional data types to the target table, while improving the flexibility of data export and the readability of the target table, thereby improving the convenience of subsequent data use and processing.

[0077] Figure 3 This is another flow chart of the data processing method provided by an embodiment of the present invention. This embodiment adds a data export step based on the above embodiment. Figure 3 As shown, the method includes:

[0078] S310: Acquire a first business data combination from a first data source and a second business data combination from a second data source within a predetermined time period.

[0079] S311. For each business data in the first business data combination and the second business data combination, if there is a missing field in the current business data and the missing field is a non-key field, the missing non-key field and the corresponding default data are added to the corresponding position of the current business data to obtain the current field missing processing result.

[0080] Taking managed business data as an example, it involves many fields such as account information, resource transfer information, resource project information, etc. It is possible to determine whether each field is a key field based on different types of fields.

[0081] The prerequisite for using the feature nearest neighbor filter to perform data integration on the first service data combination and the second service data combination is that both the first service data combination and the second service data combination must include the predetermined field combination. Therefore, before performing data integration on the two, it is necessary to detect whether both include the predetermined field combination.

[0082] If there is a missing field in the current business data, and the missing field is a key field, the corresponding field information is removed from the current business data to update the current business data, and the updated current business data is stored in a predetermined location. This processing method can reduce the impact of missing data on the overall data, thereby improving the accuracy of subsequent data integration processing; if there is no missing key field in the current business data, but there is a missing non-key field, the missing non-key field and the corresponding default data are added to the corresponding position of the current business data to obtain the current field missing processing result. This operation can make the first business data include complete data records, thereby reducing the difficulty of data integration processing.

[0083] S312. If it is determined based on the regularized expression that there is no abnormality in the current field missing processing result, the current field missing processing result is transmitted to a predetermined location. Otherwise, the abnormal field information in the current field missing processing result is removed to update the current field missing processing result, and the updated current field missing processing result is stored in a predetermined location.

[0084] Determine whether there are any abnormal fields in the current field missing processing result based on a regular expression. If no abnormal fields exist, the current field missing processing result is optionally transferred to a predetermined location based on a predetermined data query class. If there are abnormal fields, the abnormal field information in the current field missing processing result is removed to update the current field missing processing result, and the updated current field missing processing result is added to the predetermined location. Abnormal fields can include invalid fields, error fields, etc.

[0085] Among them, the predetermined data query class enables the database, the middle server interface and the terminal to transfer data in the form of objects, and each field of each business data in the database corresponds one-to-one to the variable in the predetermined data query class.

[0086] Among them, a regular expression is a special text string used to match string patterns. In data validation, it can be used to determine whether data conforms to a predetermined format, identify abnormal data that does not conform to the expected pattern, and quickly scan large amounts of data to find potential errors.

[0087] S320 : Integrate the first business data combination and the second business data combination in the predetermined storage location based on a feature nearest neighbor filter to obtain a target data set.

[0088] It is understood that after executing S311 and S312 on each of the business data in the first business data combination and the second business data combination, the processed first business data combination and the second business data combination are stored in the predetermined storage location. Data integration is performed on the processed first business data combination and the second business data combination based on the feature nearest neighbor filter to obtain a target data set.

[0089] The technical solution provided by the embodiment of the present invention ensures that each business data conforms to the predetermined data format through methods such as key field missing judgment, non-key field missing judgment and corresponding processing, and abnormal field identification and processing, thereby improving the accuracy of data integration, that is, improving the accuracy of the target data set.

[0090] Figure 4 Schematic diagram of the structure of the data processing device provided by the embodiment of the present invention. Figure 4 As shown, the device includes:

[0091] A data acquisition module 31 is configured to acquire a first business data combination from a first data source and a second business data combination from a second data source within a predetermined time period;

[0092] The integration module 32 is configured to integrate the first service data combination and the second service data combination based on a feature nearest neighbor filter to obtain a target data set.

[0093] In one embodiment, the integration module 32 includes:

[0094] a data adding unit, configured to, for each first business data in the first business data combination, add the second business data to the target data set if the second business data combination contains second business data that meets a similarity condition with the current first business data;

[0095] an offset unit, configured to determine, if the second business data combination does not contain any second business data that meets the similarity condition with the current first business data, the sum of similarity offsets of other first business data in the first business data combination relative to the corresponding second business data in the second business data combination, and determine, based on a nearest neighbor search algorithm, the local closest distance between the second business data combination and the current first business data;

[0096] An intermediate unit, configured to use the sum of the similarity offsets and the sum of the local closest distances as a current intermediate variable;

[0097] an updating unit, configured to use the current intermediate variable, the sum of the similarity offset, and the sum of the local closest distance as the updated current intermediate variable if the current first business data belongs to a reference set but does not belong to the first business data combination, wherein the reference set includes all fields in the first business data combination and the second business data combination;

[0098] The offset adding unit is configured to take the sum of the current target data set and the updated current intermediate variable as the updated current target data set when the updated current intermediate variable meets the boundary condition.

[0099] In one embodiment, the offset unit is used to:

[0100] Determine a total similarity offset between each first business data in the first business data combination and the corresponding second business data in the second business data combination, and a current similarity offset of the current first business data relative to the second business data with the same identifier in the second business data;

[0101] The difference between the total similarity offset and the current similarity offset is used as the sum of similarity offsets of other first business data in the first business data combination relative to corresponding second business data in the second business data combination.

[0102] In one embodiment, Figure 5 As shown, the device further includes an export module 33, which includes:

[0103] a response unit, configured to determine at least one query subject in response to the data export request;

[0104] A table creation unit, configured to create a target table, wherein the target table includes at least one subtable, and a subject name of the at least one subtable corresponds to the at least one query subject in a one-to-one manner;

[0105] The export unit is configured to export the data in the target data set corresponding to each query subject in the at least one query subject to a subtable corresponding to the subject name in the target table, so as to update the target table.

[0106] In one embodiment, the export module 33 further includes a configuration unit, which is configured to:

[0107] Display table configuration interface;

[0108] Determine the table name of the target table according to the name configuration information received through the table configuration interface, and determine the storage path of the target table according to the storage configuration information received through the table configuration interface;

[0109] The target table is named as the table name, and the target table is stored according to the storage path.

[0110] In one embodiment, Figure 6 As shown, the device further includes a pre-processing module 34, which includes:

[0111] a missing processing unit, configured to, for each business data in the first business data combination and the second business data combination, if a field is missing in the current business data and the missing field is a non-key field, add the missing non-key field and corresponding default data to a corresponding position of the current business data to obtain a current field missing processing result;

[0112] a field exception processing unit, configured to, if it is determined based on a regularized expression that the current field missing processing result does not contain an exception, transmit the current field missing processing result to a predetermined location; otherwise, remove the abnormal field information from the current field missing processing result to update the current field missing processing result, and store the updated current field missing processing result in a predetermined location;

[0113] The integrated module 32 is used for:

[0114] Based on a feature nearest neighbor filter, the first business data combination and the second business data combination in the predetermined storage location are integrated to obtain the target data set.

[0115] The technical solution of the data processing method provided by the embodiment of the present invention is that since the first business data combination comes from the first data source and the second business data comes from the second data source, there may be differences in the fields between the first business data combination and the second business data combination. The first business data combination and the second business data combination are integrated from multiple dimensions based on the feature nearest neighbor filter, and the first business data combination and the second business data combination can be integrated into a target data set; the technical effect of integrating the first business data combination and the second business data combination into the target data set under the premise of ensuring data integrity, reliability and scalability is achieved, which is conducive to discovering hidden relationships between data and further analysis and processing.

[0116] The data processing device provided by the embodiment of the present invention can execute the data processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0117] Figure 7 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0118] like Figure 7As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0119] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0120] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the data processing method.

[0121] In some embodiments, the data processing method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data processing method in any other suitable manner (e.g., by means of firmware).

[0122] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0123] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0124] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0125] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0126] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0127] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0128] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the data processing method provided in any embodiment of the present application.

[0129] The computer program product may be implemented by writing computer program code for performing the operations of the present invention in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0130] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0131] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A data processing method, characterized in that: include: Acquire a first business data combination from a first data source and a second business data combination from a second data source within a predetermined time period; The first service data combination and the second service data combination are integrated based on a feature nearest neighbor filter to obtain a target data set.

2. The method according to claim 1, characterized in that The integrating the first service data combination and the second service data combination based on the feature nearest neighbor filter to obtain a first data set includes: For each first business data in the first business data combination, if the second business data combination contains second business data that meets the similarity condition with the current first business data, then add the second business data to the target data set; If the second business data combination does not contain second business data that meets the similarity condition with the current first business data, then determining the sum of similarity offsets of other first business data in the first business data combination relative to the corresponding second business data in the second business data combination, and determining the local closest distance between the second business data combination and the current first business data based on a nearest neighbor search algorithm; Taking the sum of the similarity offsets and the sum of the local closest distances as the current intermediate variable; If the current first business data belongs to a reference set but does not belong to the first business data combination, the current intermediate variable, the sum of the similarity offset, and the sum of the local closest distance are used as the updated current intermediate variable, and the reference set includes all fields in the first business data combination and the second business data combination; In a case where the updated current intermediate variable meets the boundary condition, the sum of the current target data set and the updated current intermediate variable is used as the updated current target data set.

3. The method according to claim 2, characterized in that The determining the sum of similarity offsets of other first business data in the first business data combination relative to corresponding second business data in the second business data combination includes: Determine a total similarity offset between each first business data in the first business data combination and the corresponding second business data in the second business data combination, and a current similarity offset of the current first business data relative to the second business data with the same identifier in the second business data; The difference between the total similarity offset and the current similarity offset is used as the sum of similarity offsets of other first business data in the first business data combination relative to corresponding second business data in the second business data combination.

4. The method according to claim 1, wherein After obtaining the first data set, the method further includes: In response to the data export request, determining at least one query subject; Creating a target table, wherein the target table includes at least one subtable, and a subject name of the at least one subtable corresponds to the at least one query subject in a one-to-one manner; The data in the target data set corresponding to each query subject in the at least one query subject are exported to a subtable corresponding to the subject name in the target table to update the target table.

5. The method according to claim 4, characterized in that After exporting the data in the target data set corresponding to each query subject in the at least one query subject to the subtable corresponding to the subject name in the target table, the method further includes: Display table configuration interface; Determine the table name of the target table according to the name configuration information received through the table configuration interface, and determine the storage path of the target table according to the storage configuration information received through the table configuration interface; The target table is named as the table name, and the target table is stored according to the storage path.

6. The method according to claim 1, characterized in that Before integrating the first service data combination and the second service data combination based on the feature nearest neighbor filter to obtain the first data set, the method further includes: For each business data in the first business data combination and the second business data combination, if a field is missing in the current business data and the missing field is a non-key field, the missing non-key field and the corresponding default data are added to the corresponding position of the current business data to obtain a current field missing processing result; If it is determined based on the regularized expression that the current field missing processing result does not have any abnormality, the current field missing processing result is transmitted to a predetermined location; otherwise, the abnormal field information in the current field missing processing result is removed to update the current field missing processing result, and the updated current field missing processing result is stored in a predetermined location; The integrating the first service data combination and the second service data combination based on the feature nearest neighbor filter to obtain a first data set includes: Based on a feature nearest neighbor filter, the first business data combination and the second business data combination in the predetermined storage location are integrated to obtain the target data set.

7. A data processing device, characterized in that: include: A data acquisition module, configured to acquire a first business data combination from a first data source and a second business data combination from a second data source within a predetermined time period; An integration module is used to integrate the first business data combination and the second business data combination based on a feature nearest neighbor filter to obtain a target data set.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the data processing method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the data processing method according to any one of claims 1 to 7 when executed.

10. A computer program product, characterized in that The computer program product comprises a computer program, which implements the data processing method according to any one of claims 1 to 7 when executed by a processor.