Primary key determination method and device and computer readable storage medium

By determining the unified primary key in streaming data processing, the integration problem caused by the complexity of streaming data primary keys is solved, and data integration and query performance are improved.

CN120021230APending Publication Date: 2025-05-20HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410088939.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-01-22
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The primary key of streaming data is complex, resulting in the inability to integrate.

Method used

Provides a primary key determination method, by obtaining multiple data records, sampling and determining possible candidate keys, selecting target candidate keys, and generating a unified primary key for each data record.

Benefits of technology

The integration of streaming data in complex primary key situations is realized, and the efficiency and query performance of the data integration system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120021230A_ABST
    Figure CN120021230A_ABST
Patent Text Reader

Abstract

The invention provides a primary key determination method and device and a computer readable storage medium, and the method comprises the steps that a primary key generation system obtains a plurality of data records (including a plurality of same fields) in the same streaming data or different streaming data, and samples the plurality of data records, fields and / or field combinations with repeated values in the sampled data records are determined as non-candidate keys, then a plurality of possible candidate keys (the remaining elements except the non-candidate keys in a complete set are determined, and the complete set comprises a plurality of fields and combinations of at least two fields in the plurality of fields) are determined, whether the plurality of possible candidate keys can serve as candidate keys or not is judged, and if yes, the candidate keys are determined. The method comprises the following steps: selecting a target candidate key from a plurality of possible candidate keys under the condition that the possible candidate keys can be used as candidate keys, and finally generating a value of a unified primary key of each data record based on a value corresponding to the target candidate key in each data record, so that a plurality of data records can be subsequently integrated based on the unified primary key, i.e., stream data is integrated.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of the Chinese patent application filed with the State Intellectual Property Office of China on November 17, 2023, with application number 202311541738.4, and the priority of the Chinese patent application entitled "Streaming Data Processing Method, Device and Computer-readable Storage Medium", all of which are incorporated by reference in this application. Technical Field

[0002] The present application relates to the field of data processing technology, and in particular to a primary key determination method, device and computer-readable storage medium. Background Technology

[0003] Streaming data refers to data generated in a continuous time series. This type of data has a strong timeliness, reflects changes in business status, and contains a large amount of real-time information and event records, such as meteorological observation data records collected in real time by meteorological observation stations, employee punch-in records collected in real time by the company's punch-in system, user login records collected in real time by the company's service system, and commodity transaction records collected in real time by the company's transaction system. Real-time integrated processing of streaming data can effectively help make business decisions.

[0004] However, the primary key situation of streaming data is very complicated. For example, streaming data has no primary key, or some have primary keys and some do not, or the primary keys are different and cannot be integrated. SUMMARY OF THE INVENTION

[0005] This application provides a primary key determination method, device and computer-readable storage medium, which can realize streaming data integration with very complex primary key situations.

[0006] In a first aspect, a primary key determination method is provided, the method comprising the following steps:

[0007] The primary key generation system obtains multiple data records, and the multiple data records include the same multiple fields; samples the multiple data records, and determines the fields and / or field combinations with duplicate values ​​in the sampled data records as the first non-candidate key; determines the remaining elements in the full set except the first non-candidate key as multiple first possible candidate keys, wherein the full set includes multiple fields and a combination of at least two of the multiple fields; when it is determined based on the multiple data records that multiple first possible candidate keys can all be used as candidate keys, selects a first target candidate key from the multiple first possible candidate keys, or when it is determined based on the multiple data records that at least one of the multiple first possible candidate keys can be used as a candidate key, selects a first target candidate key from at least one first possible candidate key; based on the value of the corresponding first target candidate key in each data record, generates the value of the unified primary key for each data record.

[0008] The above multiple data records belong to the same data stream (also called streaming data), or belong to different data streams.

[0009] In some possible implementations, the above method further includes the following steps:

[0010] The primary key generation system adds the value of the unified primary key of each data record to each data record and sends each data record to the data integration system for integration.

[0011] The above scheme can analyze multiple data records that include the same multiple fields in the streaming data, determine possible candidate keys, and select the target candidate key after verifying the possible candidate keys. Then, based on the values ​​of the corresponding target candidate keys in the multiple data records, a unified primary key value is generated for the multiple data records. In this way, the subsequent data integration system can integrate multiple data records based on the unified primary key of multiple data records, thereby realizing the integration of streaming data.

[0012] In some possible implementations, none of the above multiple data records carry a primary key, or the multiple data records carry different primary keys, or some of the multiple data records carry a primary key and the rest do not carry a primary key.

[0013] In some possible implementations, the above method further includes the following steps:

[0014] When the primary key generation system determines that multiple first possible candidate keys cannot all be candidate keys based on multiple data records, it continues to sample multiple data records, and determines the fields and / or field combinations with duplicate values ​​in the data records that will continue to be sampled as second non-candidate keys; it determines the remaining elements in the full set except the first non-candidate key and the second non-candidate key as multiple second possible candidate keys; when it is determined that multiple second possible candidate keys can all be candidate keys based on multiple data records, it selects a second target candidate key from the multiple second possible candidate keys; based on the value of the corresponding second target candidate key in each data record, it generates the value of the unified primary key for each data record.

[0015] By analogy based on the above scheme, when the primary key generation system determines that multiple second possible candidate keys cannot all be candidate keys based on multiple data records, it can continue to sample multiple data records, and determine the fields and / or field combinations with repeated values ​​in the data records that will continue to be sampled as third non-candidate keys, ..., and generate a unified primary key value for each data record.

[0016] It can be seen that the above scheme can filter out all non-candidate keys from the entire set, leaving all candidate keys, and select appropriate candidate keys from the remaining candidate keys, and generate a unified primary key value for each data record based on the appropriate candidate keys.

[0017] In some possible implementations, the primary key generation system can select the first target candidate key from multiple first possible candidate keys in the following way: use the first possible candidate key with the most uniform data distribution among the multiple first possible candidate keys as the first target candidate key.

[0018] In a database, the primary key is usually used for indexing to accelerate data retrieval. If the values of the primary key are evenly distributed, the indexing effect will be better and the query speed will be faster.

[0019] In the above solution, since the data distribution of the first target candidate key selected from multiple first possible candidate keys is uniform, it can be understood that this can make the values of the unified primary keys generated for multiple data records based on the value of the first target candidate key also relatively uniform. Then, when the data integration system uses the values of the unified primary keys as the indexes for multiple data records, the indexing effect will be better and the query speed will be faster.

[0020] In some possible implementations, the primary key generation system can hash the value corresponding to the first target candidate key in each data record and use the obtained hash value as the value of the unified primary key for each data record.

[0021] Optionally, the primary key generation system can also use the value corresponding to the first target candidate key in each data record as the value of the unified primary key for each data record.

[0022] It can be understood that when the primary key generation system hashes the value corresponding to the first target candidate key in each data record and uses the obtained hash value as the value of the unified primary key for each data record, compared with using the value corresponding to the first target candidate key in each data record as the value of the unified primary key for each data record, since the hash function is a one-way function, that is, the original value cannot be deduced from the hash value, therefore, it can avoid the direct exposure of the value corresponding to the first target candidate key in each data record and improve the security of the data.

[0023] In addition, since hash functions can generally map the input data evenly into the output space, this can ensure that the generated hash values are relatively evenly distributed in the unified primary key. The evenly distributed primary key helps to improve the data query performance and index efficiency. Moreover, since the hash values generated by hash functions usually have a fixed length and are not affected by the length of the input data, this can ensure that the values of the unified primary key occupy a fixed space when stored and indexed, improving the storage efficiency and query performance of the database. Further, since the hash value is calculated based on the value of the corresponding first target candidate key in each data record and is independent of the value of the corresponding first target candidate key in each data record itself, this means that even if the value of the corresponding first target candidate key in each data record changes subsequently, the generated hash value will not be affected. Therefore, the value of the primary key field can remain unchanged, which is very useful for data update and maintenance and can avoid primary key conflicts and data consistency problems due to changes in the original data. In short, using hash values as the values of the unified primary key has advantages such as data protection, uniqueness guarantee, even distribution, fixed length, and independence compared to directly using the original data as the values of the unified primary key. These advantages can improve the security, performance, and maintainability of the database.

[0024] In some possible implementation manners, the above method further includes the following steps:

[0025] When the primary key generation system receives a new data record, it generates the value of the unified primary key of the new data record based on the value of the corresponding first target candidate key in the new data record, where the new data record also includes multiple fields.

[0026] In some possible implementation manners, when multiple data records belong to different data streams, the different data streams originate from the same source or different sources.

[0027] It can be seen that the primary key determination method provided by this application can generate a unified primary key for homologous streaming data and can also generate a unified primary key for cross-source streaming data.

[0028] In a second aspect, a primary key determination device is provided, and the device includes:

[0029] An acquisition module, configured to acquire multiple data records, where the multiple data records include the same multiple fields;

[0030] A sampling module, configured to sample the multiple data records;

[0031] A processing module, configured to determine the fields and / or field combinations with duplicate values in the sampled data records as the first non-candidate keys;

[0032] The processing module is configured to determine, as a plurality of first possible candidate keys, the elements remaining in the universal set except for the first non-candidate key, where the universal set includes the plurality of fields and combinations of at least two of the plurality of fields;

[0033] The processing module is further configured to, when it is determined based on the plurality of data records that all of the plurality of first possible candidate keys can serve as candidate keys, select a first target candidate key from the plurality of first possible candidate keys; or, when it is determined based on the plurality of data records that at least one of the plurality of first possible candidate keys can serve as a candidate key, select a first target candidate key from the at least one first possible candidate key;

[0034] The processing module is further configured to generate a value of a unified primary key for each of the data records based on the data corresponding to the first target candidate key in each of the data records.

[0035] In some possible implementation manners, none of the plurality of data records carry a primary key, or the plurality of data records carry different primary keys, or some of the plurality of data records carry a primary key while the remaining ones do not carry a primary key.

[0036] In some possible implementation manners, the sampling module is further configured to, when it is determined based on the plurality of data records that all of the plurality of first possible candidate keys cannot serve as candidate keys, continue to sample the plurality of data records;

[0037] The processing module is further configured to determine, as a second non-candidate key, a field and / or a combination of fields with duplicate values in the data records obtained by continuous sampling;

[0038] The processing module is further configured to determine, as a plurality of second possible candidate keys, the elements remaining in the universal set except for the first non-candidate key and the second non-candidate key;

[0039] The processing module is further configured to, when it is determined based on the plurality of data records that all of the plurality of second possible candidate keys can serve as the candidate keys, select a second target candidate key from the plurality of second possible candidate keys;

[0040] The processing module is further configured to generate a value of a unified primary key for each of the data records based on the value corresponding to the second target candidate key in each of the data records.

[0041] In some possible implementation manners, the processing module is configured to use, as the first target candidate key, the first possible candidate key with the most uniform data distribution among the plurality of first possible candidate keys.

[0042] In some possible implementation manners, the processing module is configured to hash the data corresponding to the first target candidate key in each data record, and use the obtained hash value as the value of the unified primary key of each data record.

[0043] In some possible implementation manners, the obtaining module is further configured to obtain new data records, where the new data records also include the multiple fields; the processing module is further configured to generate the value of the unified primary key of the new data records based on the values corresponding to the first target candidate key in the new data records.

[0044] In some possible implementation manners, the apparatus further includes: a sending module; the processing module is further configured to add the value of the unified primary key of each data record to each data record; the sending module is configured to send each data record for integration to a data integration system.

[0045] In some possible implementation manners, the multiple data records belong to the same data stream, or belong to different data streams.

[0046] In some possible implementation manners, in the case where the multiple data records belong to different data streams, the different data streams are from the same source or different sources.

[0047] In a third aspect, a computing device is provided, which includes a processor and a memory. The memory is used to store instructions, and the processor is used to execute the instructions so that the computing device implements the method described in the first aspect and any implementation manner of the first aspect.

[0048] In a fourth aspect, a computer-readable storage medium is provided, in which instructions are stored. When the instructions are run by a computing device or a computing device cluster, the method described in the first aspect and any implementation manner of the first aspect is implemented.

[0049] In a fifth aspect, a computing device cluster is provided, which includes at least one computing device. Each computing device in the at least one computing device includes a processor and a memory. The processor of the at least one computing device is used to execute the instructions stored in the memory of the at least one computing device so that the computing device cluster implements the method described in the first aspect and any implementation manner of the first aspect.

[0050] In a sixth aspect, a computer program product containing instructions is provided. The computer program product includes instructions that can run on a computing device or be stored in software or a program product in any available medium. When the computer program product runs on a computing device or a computing device cluster, the computing device or the computing device cluster is caused to execute the method described in the first aspect and any implementation manner of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 FIG. 1 is a schematic structural diagram of a meteorological observation data integration system according to an embodiment of the present application;

[0052] Figure 2 FIG. 2 is a schematic structural diagram of a primary key generation system according to an embodiment of the present application;

[0053] Figure 3 FIG. 3 is a schematic structural diagram of another primary key generation system according to an embodiment of the present application;

[0054] Figure 4 FIG. 4 is a schematic diagram of a streaming data integration process according to an embodiment of the present application;

[0055] Figure 5 FIG. 5 is a schematic flow diagram of a primary key generation system generating a unified primary key for a data record set according to an embodiment of the present application;

[0056] Figure 6 FIG. 6 is a schematic diagram of selecting a first target candidate key with uniform data distribution through a histogram according to an embodiment of the present application;

[0057] Figure 7 FIG. 7 is a schematic flow diagram of another primary key generation system generating a unified primary key for a data record set according to an embodiment of the present application;

[0058] Figure 8 FIG. 8 is a schematic diagram of another streaming data integration process according to an embodiment of the present application;

[0059] Figure 9 FIG. 9 is a schematic diagram of a streaming data integration process according to an embodiment of the present application;

[0060] Figure 10 FIG. 10 is a schematic structural diagram of a primary key determination device according to an embodiment of the present application;

[0061] Figure 11 FIG. 11 is a schematic structural diagram of a computing device according to an embodiment of the present application;

[0062] Figure 12 FIG. 12 is a schematic structural diagram of a computing device cluster according to an embodiment of the present application;

[0063] Figure 13 FIG. 13 is a schematic structural diagram of another computing device cluster according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] First, an application scenario related to an embodiment of the present application is introduced.

[0065] Taking the application scenario of meteorological observation data integration as an example, such as Figure 1 As shown, the meteorological observation data integration scenario includes multiple meteorological observation stations 10 and a data integration system 20.

[0066] Multiple meteorological observation stations 10 are located in different regions and are used to periodically collect meteorological observation data such as atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, rainfall, and ultraviolet index in the regions where they are located, and transmit the meteorological observation data to the data integration system 20 as streaming data in real time. If it is necessary to persistently store the streaming data or perform correlation operations, a unique primary key value can be assigned to each data record of the streaming data to facilitate subsequent processing and analysis of the streaming data. The primary key can be data unrelated to the streaming data. For example, the time when the streaming data is collected (i.e., the collection time), or it can be data related to the streaming data, such as atmospheric temperature, etc. Since the streaming data is generated by multiple different meteorological observation stations 10, the units to which different meteorological observation stations belong, the collection devices used, the communication methods, etc. may not be the same. Therefore, the situation of the primary keys of the streaming data is very complex. The situation of the primary keys of the streaming data can be divided into the following situations:

[0067] Situation ①: None of the streaming data has a primary key.

[0068] Situation ②: All of the streaming data has a primary key, but since the streaming data comes from different sources (i.e., different meteorological observation stations 10), their primary keys may be different. For example, the primary key carried by the streaming data of meteorological observation station 101 is the time when the streaming data is collected, while the primary keys carried by the streaming data of meteorological observation stations 102 - 108 are the time when the streaming data is collected and the collected atmospheric temperature.

[0069] Situation ③: Some of the streaming data does not have a primary key, and the remaining part has a primary key. For example, the streaming data of meteorological observation station 102 does not have a primary key, while the streaming data of meteorological observation stations 102 - 108 all have a primary key.

[0070] The data integration system 20 is used to integrate the streaming data received in real time from multiple meteorological observation stations 10. The data integration system 20 can be a data lake, a data middle platform, or a real-time data warehouse, etc. Integration means integrating the streaming data from a certain data source or from different data sources but belonging to the same business type into a unified data storage (such as a unified data table or a unified partition) for operations such as analysis, query, and application development. The purpose of integration is to solve problems such as dispersion, redundancy, and inconsistency of streaming data, so that the streaming data can be more conveniently accessed and utilized. Specifically, the data integration system 20 can perform integration based on the primary keys of the streaming data, and integrate the streaming data with the same primary key into a unified data storage.

[0071] It should be understood that Figure 1 merely as an example of the meteorological observation data integration system, the number of meteorological observation stations 10 and data integration systems 20 can be one or any number, and the present application does not make specific limitations.

[0072] It can be understood that in addition to the above meteorological observation data integration scenario, there can be other application scenarios. For example, in the enterprise employee clock-in record integration scenario, the streaming data in this scenario includes employee ID number, employee name, clock-in time, clock-in location, clock-in method, clock-in type, overtime information, abnormal situations, etc. Among them, the clock-in method can be swiping a card, fingerprint recognition, facial recognition, mobile application, etc., the clock-in type can be work clock-in, off-duty clock-in, overtime clock-in, etc., the overtime information can be whether overtime and the start time and end time of overtime, and the abnormal situations can be being late, leaving early, etc.; for example, in the user login record integration scenario of each enterprise service system, the streaming data in this scenario includes user name, password, name, role / permission, email address, mobile phone number, department / organization, login time, login internet protocol address (abbreviated as IP address), etc.; and for another example, in the commodity transaction record integration scenario of each enterprise transaction system, the streaming data in this scenario includes that the commodity transaction record can include the purchase account name, purchased commodity type, purchased commodity quantity, amount paid, order number, order creation time, order payment time, order delivery time, and order receipt time, etc.

[0073] However, due to the complex primary key situation of the streaming data, the data integration system 20 cannot perform integration.

[0074] To address the above problems, the present application provides a primary key generation system, a primary key determination method, a device, etc. It can, in the case where multiple streaming data have no primary key, or in the case where some of the multiple streaming data have a primary key and some do not, or in the case where the primary keys of the multiple streaming data are different, generate a unified primary key for the multiple streaming data by analyzing the multiple streaming data, so that the data integration system can integrate the multiple streaming data based on the unified primary key.

[0075] Next, the primary key generation system, the primary key determination method, the device, etc. provided by the application will be introduced in detail with reference to the corresponding drawings respectively.

[0076] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the architecture of a primary key generation system provided by an embodiment of the present application. As Figure 2 shown, the architecture includes a client 100, a primary key generation system 200, and a data integration system 300.

[0077] The client 100 can be deployed at the source end (such asFigure 1 In the meteorological observation data integration scenario shown, the meteorological observation station 10, the employee clock-in record collection device (such as a clock-in machine) in the enterprise employee clock-in record integration scenario, the user login record collection device in the user login record integration scenario of each enterprise service system, the commodity transaction record collection device in the commodity transaction record integration scenario of each enterprise transaction system, etc.

[0078] In a specific implementation, the client 100 can be used to implement human-computer interaction and can be a device controlled by a user (such as the above-mentioned meteorological observation station 10, employee clock-in record collection device, user login record collection device, commodity transaction record collection device) or software or application programs running on a computing device, such as a personal computer client, or can also be a world wide web client accessed based on a browser, or can also be an application (APP) client running on a mobile terminal. This application does not make specific limitations.

[0079] Optionally, the client 100 can be a separate client specifically used to implement the generation of the streaming data primary key, such as a primary key generation tool, a primary key generation application, etc. Or, the client 100 can also be a primary key generation function module or plug-in within a comprehensive software, such as a streaming data primary key generation module within commonly used enterprise data integration software. This application does not make specific limitations.

[0080] Optionally, the client 100 can also be a client of a cloud platform, such as a console of the cloud platform. Specifically, it can be a web-based console or a console based on an application programming interface (API). This application does not make specific limitations. This console can provide a cloud service for generating the streaming data primary key to users, and users can obtain the usage right of the primary key generation system 200 provided by this application by purchasing the cloud service.

[0081] The primary key generation system 200 can be deployed on a computing device or a cluster of computing devices. Among them, the computing device includes a bare metal server (BMS), a virtual machine (VM), a container, or an edge computing device. A BMS refers to a general physical server, for example, an ARM server or an X86 server; a virtual machine refers to a complete computer system with the functions of a complete hardware system simulated by software and running in a completely isolated environment. All the work that can be completed on a physical computer can be achieved in a virtual machine. When creating a virtual machine in a computing device, part of the hard disk and memory capacity of the physical machine needs to be used as the hard disk and memory capacity of the virtual machine. Each virtual machine has an independent basic input / output system (BIOS), hard disk, and operating system, and can be operated like a physical machine; a container is a portable software unit that can combine an application and all its dependencies into a software package, and this software package is not restricted by the underlying host operating system, so there is no need to build a complex environment anymore, simplifying the process from application development to deployment; an edge computing device refers to a device that is closer to the data source and end user and has the characteristics of low latency and high bandwidth, such as an intelligent router, an edge server, etc. The cluster of computing devices can include multiple of the above-mentioned computing devices, such as a data center, which is not specifically limited in this application.

[0082] The data integration system 300 can be deployed on a computing device or a cluster of computing devices. The introduction of the computing device or the cluster of computing devices refers to the above text.

[0083] Optionally, the primary key generation system 200 can be deployed in the same computing device as the data integration system 300, or in different computing devices of the same cluster of computing devices, or in a computing device and a storage array of the same cluster of computing devices, or in different clusters of computing devices, which is not specifically limited in this application.

[0084] Optionally, the primary key generation system 200 can be deployed in the same or different clusters of computing devices as the client 100. For example, the client 100 is deployed on a computing device in the first cluster of computing devices, and the primary key generation system 200 is deployed on a computing device in the second cluster of computing devices; or, the client 100 and the primary key generation system 200 are deployed in the same cluster of computing devices. It should be understood that the above examples are for illustration and are not specifically limited in this application.

[0085] In some other possible implementation manners, all functions of the above primary key generation system 200 can also be implemented by the source end that generates streaming data. For example, the source end implements the above primary key generation service to generate a unified primary key for the streaming data generated by itself, or the source end implements the above primary key generation service to generate a unified primary key for the streaming data generated by other source ends.

[0086] There is a communication connection between the client 100, the primary key generation system 200, and the data integration system 300, which can be a wired connection or a wireless connection, and the present application does not make specific limitations.

[0087] It should be understood that Figure 2 The shown architecture is only for example. For example, in actual applications, it includes network devices for forwarding communication data between the primary key generation system 200, the client 100, and the data integration system 300. The number of clients 100 / data integration systems 300 that establish communication connections with the primary key generation system 200 can be one or more, and the present application does not make specific limitations. The number of primary key generation systems 200 can also be one or more, and the present application does not make specific limitations.

[0088] Please refer to Figure 3 , Figure 3 is a schematic diagram of the architecture of another primary key generation system provided by an embodiment of the present application. As shown in Figure 3 , this architecture includes a client 100, a primary key generation system 200, and a data integration system 300.

[0089] Combined with Figure 2 it can be seen that Figure 3 The difference between the shown architecture and Figure 2 the shown architecture is that the client 100 is deployed independently of the source end. Specifically, the client 100 can be deployed on a computing device or a computing device cluster. The introduction of the computing device or the computing device cluster refers to the above.

[0090] Regarding Figure 3 the same parts between the shown architecture and Figure 2 the shown architecture, please refer to Figure 2 the relevant descriptions in, and details will not be elaborated here.

[0091] It should be understood that Figure 3 The shown architecture is only for example. For example, in actual applications, it includes network devices for forwarding communication data between the primary key generation system 200, the client 100, and the data integration system 300. The number of clients 100 / data integration systems 300 that establish communication connections with the primary key generation system 200 can be one or more, and the present application does not make specific limitations. The number of primary key generation systems 200 can also be one or more, and the present application does not make specific limitations.

[0092] To facilitate a clearer understanding Figure 2 The architecture shown generates a unified primary key for multiple streams of data and integrates multiple streams of data based on the unified primary key. The specific process is as follows. Below, in conjunction with Figure 4 a schematic diagram of a process for integrating streaming data provided by an embodiment of the present application shown, a more detailed introduction will be given. It should be noted that in Figure 4 , it is described by taking Figure 2 including two source ends, with each source end deploying a client respectively as an example. As shown in Figure 4 , the method includes the following steps:

[0093] S401: The first client obtains the first stream of data.

[0094] The first stream of data is the data stream obtained by the first client within a period of time. Among them, the first stream of data can be meteorological observation data, employee clock-in records, user login records, commodity transaction records, etc. The present application does not limit the first stream of data. The first stream of data includes multiple fields. Moreover, when the first stream of data is meteorological observation data, the fields of the first stream of data include one or more of atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, and rainfall, etc. When the first stream of data is employee clock-in records, the fields of the first stream of data include one or more of employee ID number, employee name, clock-in time, clock-in location, clock-in method, clock-in type, overtime information, abnormal conditions, etc. When the first stream of data is employee clock-in records, the fields of the first stream of data include one or more of purchase account name, purchased commodity type, purchased commodity quantity, amount paid, order number, order creation time, order payment time, order delivery time, and order receipt time, etc. When the first stream of data is user login records, the fields of the first stream of data include one or more of username, password, name, role / permission, email address, mobile phone number, department / organization, login time, login internet protocol address (abbreviated as IP address), etc.

[0095] The first stream of data includes multiple data records. For example, when the first stream of data is meteorological observation data, the multiple data records are one or more of atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, and rainfall, etc. collected by the first client at different times.

[0096] S402: The first client sends the first stream of data to the primary key generation system. Correspondingly, the primary key generation system receives the first stream of data sent by the first client.

[0097] S403: The second client obtains the second stream of data.

[0098] The second streaming data is the data stream obtained by the second client within a period of time. Among them, the second streaming data can be meteorological observation data, employee clock-in records, user login records, commodity transaction records, etc., and the present application does not limit the second streaming data. The second streaming data includes multiple fields. When the second streaming data is meteorological observation data, the fields of the second streaming data include one or more of atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, and rainfall, etc. When the second streaming data is employee clock-in records, the fields of the second streaming data include one or more of employee ID, employee name, clock-in time, clock-in location, clock-in method, clock-in type, overtime information, abnormal conditions, etc.; when the second streaming data is employee clock-in records, the fields of the second streaming data include one or more of purchase account name, purchased commodity type, quantity of purchased commodities, amount paid, order number, order creation time, order payment time, order delivery time, and order receipt time, etc.; when the second streaming data is user login records, the fields of the second streaming data include one or more of username, password, name, role / permission, email address, mobile phone number, department / organization, login time, login internet protocol address (abbreviated as IP address), etc.

[0099] The second streaming data includes multiple data records. For example, when the second streaming data is meteorological observation data, the multiple data records are one or more of atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, and rainfall, etc. collected by the second client at different times.

[0100] The fields of the second streaming data and the first streaming data are the same. That is to say, the first streaming data and the second streaming data are data of the same business type. For example, assume that the first streaming data and the second streaming data both include the following fields: purchase account name, purchased commodity type, quantity of purchased commodities, amount paid, order number, order creation time, order payment time, order delivery time, and order receipt time. Then, the first streaming data and the second streaming data are both commodity transaction records. Another example is that the first streaming data and the second streaming data both include the following fields: atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, rainfall, and ultraviolet index. Then, the first streaming data and the second streaming data are both meteorological observation information. However, the order of the fields can be the same or different. For example, for the first streaming data and the second streaming data, both include fields such as atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, and rainfall. However, the order of the fields in the first streaming data is atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, and rainfall, while the order of the multiple fields in the second streaming data is atmospheric humidity, atmospheric temperature, wind speed, air pressure, wind direction, and rainfall.

[0101] S404: The second client sends second streaming data to the primary key generation system. Correspondingly, the primary key generation system receives the first streaming data sent by the second client.

[0102] S405: The primary key generation system generates a unified primary key for multiple data records in the first streaming data and multiple data records in the second streaming data.

[0103] The unified primary key includes multiple values, and each value is a unique identifier for multiple data records in the first streaming data and multiple data records in the second streaming data. Each value of the unified primary key is generated based on the corresponding field values in the same unique fields or combinations of fields in the first streaming data and the second streaming data. Among them, the way that each value of the unified primary key is generated based on the corresponding field values in the same fields or combinations of fields in the first streaming data and the second streaming data is: The value in the unified primary key can be the value of the unique fields or combinations of fields in the first streaming data and the second streaming data, or is a hash value obtained by hashing the value of the unique fields or combinations of fields in the first streaming data and the second streaming data. The meaning of the unique fields or combinations of fields is: Taking field A in the first streaming data and the second streaming data as an example, there are no duplicate values in the values of field A in the first streaming data and the second streaming data.

[0104] For the specific implementation process of S405, please refer to the following Figure 5 、 Figure 7 related description.

[0105] S406: The primary key generation system adds the unified primary key generated for each data record to each data record.

[0106] Adding the unified primary key to each data record means adding a unified primary key field carrying the corresponding key value to each data record.

[0107] Furthermore, the unified primary key can be added in front of the first field of multiple fields in each data record, or added behind the last field of multiple fields, or added between any two fields of multiple fields. This application does not make specific limitations on this.

[0108] S407: The primary key generation system sends each data record carrying the unified primary key to the data integration system. Correspondingly, the data integration system receives each data record carrying the unified primary key sent by the primary key generation system.

[0109] S408: The data integration system integrates each data record based on the unified primary key carried by each data record.

[0110] Taking the unified primary key field carried by each data record as A, when the data integration system receives the first data record among multiple data records in the first streaming data and multiple data records in the second streaming data, it can establish a data table (or partition) corresponding to the primary key field A based on the primary key field A, and store the first data record into the data table (or partition) corresponding to the primary key field A. Subsequently, when the data integration system receives other data records carrying the primary key field A, it can find the corresponding data table (or partition) based on the primary key field A carried by the other data records, and then store the other data records into the data table (or partition) corresponding to the primary key field A.

[0111] In a specific embodiment of the present application, when the data integration system subsequently receives other data records, it can also compare the primary key field and the primary key value carried by the latest received data record with the primary key field A and the primary key value carried by the already received data records. If it is determined that the primary key field and the primary key value are both the same, it is determined that the latest received data record is a duplicate data record, and it can be discarded or subjected to other processing. Otherwise, it is determined that the latest received data record is not a duplicate data record, and the data record is integrated into the data table (or partition) found corresponding to the primary key field A.

[0112] In some possible embodiments, when integrating each data record, the data integration system can also use the value of the unified primary key of each data record as the index of the data record, so that the data record can be located and searched based on the index subsequently.

[0113] Next, first in combination with Figure 5 , introduce an implementation process in which the primary key generation system generates a unified primary key for a data record set (referring to multiple data records in the first streaming data and multiple data records in the second streaming data), as Figure 5 shown, including the following steps:

[0114] S501: The primary key generation system samples the data record set, and determines the fields and / or field combinations with duplicate values in the sampled data records as the first non-candidate keys.

[0115] The first non-candidate keys are fields or field combinations that cannot be used to uniquely identify the sampled data records because there are duplicate values in the values corresponding to the first non-candidate keys in the sampled data records. It can be understood that the first non-candidate keys cannot be used to uniquely identify the sampled data records. Since the sampled data records are derived from the data record set, the first non-candidate keys cannot be used to uniquely identify the data record set either.

[0116] It can be understood that when the sampled data records are different, the first non-candidate keys determined based on the sampled data records are usually also different. The following will be described in conjunction with the meteorological observation information table shown in Table 1. It should be noted that Figure 1 The data in the meteorological observation information table shown is only for example and is not regarded as a specific limitation.

[0117] Table 1 Meteorological Observation Information Table

[0118]

[0119]

[0120] Taking the data record set including the five data records shown in Table 1 as an example:

[0121] Example 1. Suppose the data records sampled by the primary key generation system are the first row data, the second row data, and the fourth row data in Table 1. Then, by comparing the first row data, the second row data, and the fourth row data, the primary key generation system can determine that there are duplicate data "15.0" in the atmospheric temperature field, duplicate data "25.6" in the atmospheric humidity field, and duplicate data "15.0 + 25.6" in the "atmospheric temperature + atmospheric humidity" field combination. Therefore, the primary key generation system can determine that the fields with the same data are the atmospheric temperature field, the atmospheric humidity field, and the "atmospheric temperature + atmospheric humidity" field combination. Then, the primary key generation system determines the atmospheric temperature field, the atmospheric humidity field, and the "atmospheric temperature + atmospheric humidity" field combination as the three first non-candidate keys.

[0122] Example 2. Suppose the data records sampled by the primary key generation system are the first row data, the second row data, the fourth row data, and the fifth row data in Table 1. Then, by comparing the first row data, the second row data, the fourth row data, and the fifth row data, the primary key generation system can determine that there are duplicate data "14.5" in the atmospheric temperature field, duplicate data "25.6" in the atmospheric humidity field, duplicate data "0.11" in the wind speed field, duplicate data "15.0 + 25.6" in the "atmospheric temperature + atmospheric humidity" field combination, and duplicate data "14.5 + 0.11" in the "atmospheric temperature + wind speed" field combination. Therefore, the primary key generation system can determine that the fields with the same data are the atmospheric temperature field, the atmospheric humidity field, the wind speed field, the "atmospheric temperature + atmospheric humidity" field combination, and the "atmospheric temperature + wind speed" field combination. Then, the primary key generation system determines the atmospheric temperature field, the atmospheric humidity field, the wind speed field, the "atmospheric temperature + atmospheric humidity" field combination, and the "atmospheric temperature + wind speed" field combination as the five first non-candidate keys.

[0123] Based on the above examples, it can be seen that when the sampled data records are different, the first non-candidate key determined based on the sampled data records is also different.

[0124] In this application, since the data volume of the data record set is usually relatively large, such as including 5,000 or 10,000 data records, the primary key generation system samples the data record set and determines the first non-candidate key based on the sampled data records, rather than determining the first non-candidate key based on all data records, which can improve the efficiency of determining the first non-candidate key, thereby improving the efficiency of generating a unified primary key.

[0125] S502: The primary key generation system determines the remaining elements in the full set except the first non-candidate key as multiple first possible candidate keys, where the full set includes multiple fields and a combination of at least two of the multiple fields.

[0126] Continue taking the observation time, atmospheric temperature, atmospheric humidity, and wind speed as shown in Table 1 as an example, the full set includes: observation time field, atmospheric temperature field, atmospheric humidity field, wind speed field, "observation time + atmospheric temperature" field combination, "observation time + atmospheric humidity" field combination, "observation time + wind speed" field combination, "atmospheric temperature + atmospheric humidity" field combination, "atmospheric temperature + wind speed" field combination, "atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity" field combination, "observation time + atmospheric temperature + wind speed" field combination, "observation time + atmospheric humidity + wind speed" field combination, "atmospheric temperature + atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination, "atmospheric temperature + atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination.

[0127] From the relevant description in S501, it can be seen that when the sampled data records are different, the first non-candidate key determined based on the sampled data records is also different. It can be understood that when the first non-candidate key is different, the first possible candidate keys obtained by filtering the full set based on the first non-candidate key are usually also different. The following is an explanation in combination with Example 1 and Example 2 in S501.

[0128] Example (1), firstly, taking the atmospheric temperature field, atmospheric humidity field and the "atmospheric temperature + atmospheric humidity" field combination described in Example 1 in S501 as an example, the multiple first possible candidate keys determined by the primary key generation system include: observation time field, wind speed field, "observation time + atmospheric temperature" field combination, "observation time + atmospheric humidity" field combination, "observation time + wind speed" field combination, "atmospheric temperature + wind speed" field combination, "atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity" field combination, "observation time + atmospheric temperature + wind speed" field combination, "observation time + atmospheric temperature + wind speed" field combination, "observation time + atmospheric humidity + wind speed" field combination, "atmospheric temperature + atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination.

[0129] Example (2), taking the atmospheric temperature field, atmospheric humidity field, wind speed field, "atmospheric temperature + atmospheric humidity" field combination and "atmospheric temperature + wind speed" field combination described in Example 2 in S501 as an example, the multiple first possible candidate keys determined by the primary key generation system include: observation time field, "observation time + atmospheric temperature" field combination, "observation time + atmospheric humidity" field combination, "observation time + wind speed" field combination, "atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity" field combination, "observation time + atmospheric temperature + wind speed" field combination, "observation time + atmospheric humidity + wind speed" field combination, "observation time + atmospheric humidity + wind speed" field combination, "atmospheric temperature + atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination.

[0130] Based on the above examples, it can be seen that when the first non-candidate keys are different, filtering the entire set based on the first non-candidate key will also result in multiple first possible candidate keys that are different.

[0131] S503: The primary key generation system verifies whether multiple first possible candidate keys can all be used as candidate keys based on the data record set. If it is determined that multiple first possible candidate keys can all be used as candidate keys, S504 is executed. If it is determined that multiple first possible candidate keys cannot all be used as candidate keys, S505-S507 are executed.

[0132] A candidate key is a field or combination of fields that can be used to uniquely identify a data record set. In other words, there are no duplicate values ​​in the values ​​corresponding to the candidate key in the data record set. Therefore, each value corresponding to the candidate key in the data record set can uniquely identify the data record to which each value belongs.

[0133] Continuing with Table 1 as an example, it can be seen from Table 1 that there are no duplicate values in the corresponding values of the observation time field, the "observation time + atmospheric temperature" field combination, the "observation time + atmospheric humidity" field combination, and the "observation time + atmospheric temperature + atmospheric humidity" field combination. Therefore, these fields and field combinations can all be used to uniquely identify the data records in Table 1 and are all candidate keys.

[0134] As can be seen from the relevant description in S502, when the first non-candidate keys are different, filtering the entire set based on the first non-candidate keys results in different multiple first possible candidate keys. The multiple first possible candidate keys may all be candidate keys or may not all be candidate keys. Therefore, it is necessary to verify whether the multiple first possible candidate keys can all be candidate keys. When it is determined that the multiple first possible candidate keys can all be candidate keys, S504 is executed, which can ensure that the first target candidate key selected from the multiple first possible candidate keys can necessarily be used to uniquely identify the data record set. As a result, the value of the unified primary key generated for each data record based on the value of the corresponding first target candidate key in each data record in the data record set can necessarily be used to uniquely identify each data record, ensuring the accuracy of the unified primary key. When it is determined that there are non-candidate keys among the multiple first possible candidate keys, S505 and S506 are executed to sample more data records to filter out more or even all non-candidate keys from the entire set to ensure the accuracy of the subsequent generated unified primary key.

[0135] In a possible embodiment, the primary key generation system can verify whether the multiple first possible candidate keys can all be candidate keys when the number of the multiple first possible candidate keys reaches a certain threshold, and does not verify whether the multiple first possible candidate keys can all be candidate keys when the number of the multiple first possible candidate keys does not reach a certain threshold. Here, the threshold can be customized according to the actual scenario. For example, the threshold can be set to 3 or 5, etc. The present application does not make specific limitations on this.

[0136] When determining whether to verify whether the multiple first possible candidate keys can all be candidate keys, the verification process can be illustrated by taking the i-th first possible candidate key among the multiple first possible candidate keys as an example: The primary key generation system can determine whether there are duplicate values in the values of the corresponding i-th first possible candidate key in the data record set. If it is determined that there are duplicate values, it is determined that the i-th first possible candidate key cannot be a candidate key. If it is determined that there are no duplicate values, it is determined that the i-th first possible candidate key can be a candidate key.

[0137] The process of verifying the first possible candidate keys is introduced below in combination with Example (1) and Example (2) in S502.

[0138] First, taking the multiple first possible candidate keys described in Example (1) of S502 as an example, which include: the observation time field, the wind speed field, the "observation time + atmospheric temperature" field combination, the "observation time + atmospheric humidity" field combination, the "observation time + wind speed" field combination, the "atmospheric temperature + wind speed" field combination, the "atmospheric humidity + wind speed" field combination, the "observation time + atmospheric temperature + atmospheric humidity" field combination, the "observation time + atmospheric temperature + wind speed" field combination, the "observation time + atmospheric humidity + wind speed" field combination, the "atmospheric temperature + atmospheric humidity + wind speed" field combination, and the "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination. Since there are duplicate values "0.11" in the wind speed field and duplicate value "14.5 + 0.11" in the "atmospheric temperature + wind speed" field combination, while there are no duplicate values in the remaining fields and field combinations, the primary key generation system can determine that not all of the multiple first possible candidate keys can be used as candidate keys.

[0139] Taking the multiple first possible candidate keys described in Example (2) of S502 as an example, which include: the observation time field, the "observation time + atmospheric temperature" field combination, the "observation time + atmospheric humidity" field combination, the "observation time + wind speed" field combination, the "atmospheric humidity + wind speed" field combination, the "observation time + atmospheric temperature + atmospheric humidity" field combination, the "observation time + atmospheric temperature + wind speed" field combination, the "observation time + atmospheric humidity + wind speed" field combination, the "atmospheric temperature + atmospheric humidity + wind speed" field combination, and the "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination. Since there are no duplicate values in these fields and field combinations, the primary key generation system can determine that all of the multiple first possible candidate keys can be used as candidate keys.

[0140] S504: The primary key generation system selects a first target candidate key from the multiple first possible candidate keys, and generates a unified primary key value for each data record based on the value of the corresponding first target candidate key in each data record.

[0141] The ways for the primary key generation system to select a first target candidate key from the multiple first possible candidate keys at least include the following:

[0142] Possible selection method (1): Select any one of the multiple first possible candidate keys as the first target candidate key.

[0143] Possible selection method (2): Select the first possible candidate key with the most uniform data distribution among the multiple first possible candidate keys as the first target candidate key, where the data distribution of the first possible candidate key is all the data corresponding to the first possible candidate key in the data record set.

[0144] Possible selection method (3): Select the first possible candidate key with the most uniform data distribution and the fewest number of fields among multiple first possible candidate keys as the first target candidate key.

[0145] In a specific implementation, the first possible candidate key with a uniform (or the most uniform) data distribution can be selected in the following way: By plotting the frequency distribution histograms of the data of multiple first possible candidate keys and observing the data distribution in different intervals, the first possible candidate key can be found from multiple first possible candidate keys and used as the first target candidate key. Optionally, statistical metrics can also be used to judge the uniformity of the data, such as the mean, median, and mode. The mean is used to understand the overall mean of the data. If the mean of the data is not much different from the median and mode, it indicates that the data is relatively uniform. Another example is variance and standard deviation. Variance and standard deviation reflect the degree of dispersion of the data. If the variance or standard deviation is small, the data is relatively uniform. If the variance or standard deviation is large, the data is relatively non-uniform. This application does not limit the method of judging the data distribution uniformity of multiple first possible candidate keys.

[0146] Taking the judgment of the data distribution uniformity of multiple first possible candidate keys through the frequency distribution histogram as an example, assume that two first possible candidate keys correspond to the atmospheric temperature field and the atmospheric humidity field, each corresponding to 6000 values. The histogram of the atmospheric temperature field is as shown in Figure 6 (a) in, and the histogram of the atmospheric humidity field is as shown in Figure 6 (b) in. Then the primary key generation system can determine the data distribution of the atmospheric temperature field through these two histograms: there are 1000 values less than 29°C, 1500 values between 29 - 30°C, 2800 values between 30 - 31°C, and 700 values greater than 31°C. And it can determine the data distribution of the atmospheric humidity field: there are 1500 values less than 25%, 1600 values between 25% - 26%, 1800 values between 26% - 27%, and 1100 values greater than 27%. Since the data distribution of the atmospheric humidity field is more uniform than that of the atmospheric temperature field, the primary key generation system preferably selects the atmospheric humidity field as the first target candidate key.

[0147] Through the above possible selection method (2) / (3), since the data distribution of the first target candidate key is uniform, it can be understood that this can make the data distribution of the unified primary key generated based on the value of the first target candidate key relatively uniform. Since in a database, the primary key is usually used for indexing to accelerate data retrieval, if the data distribution of the primary key is uniform, the indexing effect will be better and the query speed will be faster.

[0148] It should be understood that the above possible selection methods (1) to (3) are only example selection methods for the primary key generation system to select the first target candidate key from multiple first possible candidate keys. In specific implementation, other methods can also be referred to for selection. For example, the primary key generation system provides multiple first possible candidate keys to the user, and then obtains the first possible candidate key selected by the user from the multiple first possible candidate keys as the first target candidate key. The present application does not specifically limit the selection method.

[0149] In a possible embodiment, the primary key generation system may hash the value corresponding to the first target candidate key in each data record, and use the obtained hash value as the value of the unified primary key for each data record.

[0150] In another possible embodiment, the primary key generation system may use the value corresponding to the first target candidate key in each data record as the value of the unified primary key for each data record.

[0151] It can be understood that, compared with using the value corresponding to the first target candidate key in each data record as the value of the unified primary key for each data record, when the primary key generation system hashes the value corresponding to the first target candidate key in each data record and uses the obtained hash value as the value of the unified primary key for each data record, since the hash function is a one-way function, that is, the original value cannot be deduced from the hash value, the value corresponding to the first target candidate key in each data record can be avoided from being directly exposed, thereby improving the security of the data.

[0152] In addition, since the hash function can usually map the input data evenly into the output space, this can ensure that the generated hash values are relatively evenly distributed in the unified primary key. The evenly distributed primary key helps to improve the data query performance and index efficiency. Moreover, since the hash values generated by the hash function usually have a fixed length and are not affected by the length of the input data, this can ensure that the values of the unified primary key occupy a fixed space when stored and indexed, improving the storage efficiency and query performance of the database. Further, since the hash value is calculated based on the value corresponding to the first target candidate key in each data record and is independent of the value corresponding to the first target candidate key in each data record itself, this means that even if the value corresponding to the first target candidate key in each data record changes subsequently, the generated hash value will not be affected. Therefore, the value of the primary key field can remain unchanged, which is very useful for data update and maintenance and can avoid primary key conflicts and data consistency problems caused by changes in the original data. In short, using the hash value as the value of the unified primary key has advantages such as data protection, uniqueness guarantee, even distribution, fixed length, and independence compared with directly using the original data as the value of the unified primary key. These advantages can improve the security, performance, and maintainability of the database.

[0153] S505: The primary key generation system continues to sample the data record set, and determines the fields and / or field combinations with duplicate values ​​in the data records to be sampled as the second non-candidate key.

[0154] The second non-candidate key is a field or field combination that cannot be used to uniquely identify the data records to be sampled. This is because there are duplicate values ​​in the values ​​corresponding to the second non-candidate key in the data records to be sampled. It can be understood that the second non-candidate key cannot be used to uniquely identify the data records to be sampled. Since the data records to be sampled come from the data record set, the second non-candidate key cannot be used to uniquely identify the data record set.

[0155] S506: The primary key generation system determines the remaining elements in the entire set except the first non-candidate key and the second non-candidate key as multiple second possible candidate keys.

[0156] For an introduction to the entire collection, please refer to the introduction in S502.

[0157] S507: The primary key generation system verifies whether multiple second possible candidate keys can all be used as candidate keys based on the data record set. If it is determined that multiple second possible candidate keys can all be used as candidate keys, S508 is executed. If it is determined that multiple second possible candidate keys cannot all be used as candidate keys, S509 is executed.

[0158] S508: The primary key generation system selects a second target candidate key from multiple second possible candidate keys, and generates a unified primary key value for each data record based on the value of the second target candidate key in each data record.

[0159] In a possible embodiment, the primary key generation system may hash the value corresponding to the second target candidate key in each data record, and use the obtained hash value as the value of the unified primary key of each data record.

[0160] In another possible embodiment, the primary key generation system may use the value corresponding to the second target candidate key in each data record as the value of the unified primary key of each data record.

[0161] S509: The primary key generation system continues to sample the data record set, and determines the fields and / or field combinations with duplicate values ​​in the data records that will continue to be sampled as the third non-candidate key.

[0162] The specific implementation process of S505-S509 is similar to that of S501-S505. Please refer to the relevant description of S501-S505. For the sake of brevity, no further details will be given.

[0163] It can be seen that based on Figure 5By analogy with the embodiments, the primary key generation system can filter out all non-candidate keys from the entire set, leaving all candidate keys, and select appropriate candidate keys from the remaining candidate keys, and generate a unified primary key value for each data record based on the appropriate candidate keys.

[0164] In a possible embodiment, when the primary key generation system selects the first target candidate key / second target candidate key, when obtaining new data records including the same multiple fields, it can perform hashing based on the values corresponding to the first target candidate key / second target candidate key in the new data records, and use the obtained hash value as the unified primary key value of the new data records, add it to the new data records, or use the value corresponding to the first target candidate key in the new data records as the unified primary key value of the new data records, add it to the new data records, and then send the new data records carrying the corresponding unified primary key value to the data integration system for integration.

[0165] In a possible embodiment, when the primary key generation system selects the first target candidate key / second target candidate key, it can also send a notification carrying the first target candidate key / second target candidate key to the client (including the first client and the second client), instructing the client to perform hashing on the values corresponding to the first target candidate key / second target candidate key in the new data records when obtaining new data records including the same multiple fields, use the obtained hash value as the unified primary key value of the new data records, add it to the new data records, or use the value corresponding to the first target candidate key / second target candidate key in the new data records as the unified primary key value of the new data records, add it to the new data records, and then send the new data records carrying the corresponding unified primary key value to the data integration system for integration.

[0166] Next, in combination with Figure 7 , introduce another implementation process for the primary key generation system to generate a unified primary key for a data record set (referring to multiple data records in the first streaming data and multiple data records in the second streaming data), as Figure 7 shown, including the following steps:

[0167] S701: The primary key generation system samples the data record set, and determines the fields and / or field combinations with duplicate values in the sampled data records as the first non-candidate keys.

[0168] S701 is the same as S501, please refer to the relevant description of S501.

[0169] S702: The primary key generation system determines the remaining elements in the entire set except the first non-candidate keys as multiple first possible candidate keys, where the entire set includes multiple fields and combinations of at least two fields in the multiple fields.

[0170] S702 is the same as S502. Please refer to the relevant description of S502.

[0171] S703: The primary key generation system is based on a set of data records to verify whether there is at least one first possible candidate key among multiple first possible candidate keys that can be used as a candidate key. In the case of determining that there is at least one first possible candidate key that can be used as a candidate key, S704 is executed. In the case of determining that there is no at least one first possible candidate key that can be used as a candidate key, S705 is executed.

[0172] For the introduction of candidate keys, please refer to the introduction of candidate keys in S503.

[0173] The specific process by which the primary key generation system is based on a set of data records to verify whether there is at least one first possible candidate key among multiple first possible candidate keys that can be used as a candidate key is similar to the specific process described in S503 by which the primary key generation system is based on a set of data records to verify whether multiple first possible candidate keys can all be used as candidate keys. Please refer to the relevant description in S503.

[0174] S704: The primary key generation system selects a first target candidate key from at least one first possible candidate key and generates a unified primary key value for each data record based on the value of the corresponding first target candidate key in each data record.

[0175] S704 is similar to S504. Please refer to the relevant description of S504.

[0176] S705: The primary key generation system sends a notification to the client, notifying the user of the client that a unified primary key cannot be generated.

[0177] Next, in combination with a more detailed example, Figure 2 the entire process of streaming data integration for the Figure 8 shown system will be introduced. Refer to Figure 8 which is a schematic diagram of another streaming data integration process provided by an embodiment of the present application. Figure 8 The process from left to right is as follows:

[0178] First, the clients 100 on multiple source ends transmit the data records generated by the source ends to the primary key generation system 200 in real time.

[0179] S1, the primary key generation system 200 determines whether each received data record carries a valid primary key. In the case of determining that the data record carries a valid primary key, it sends the data record carrying the valid primary key to the data integration system 300 for integration. In the case of determining that the data record does not carry a valid primary key, S2 is executed.

[0180] Specifically, the primary key generation system 200 can pre-store multiple primary key fields of the data already integrated in the data integration system 300. When receiving each data record, it determines whether the data record carries a primary key field. If it is determined that the data record does not carry a primary key field, it is determined that the data record does not carry a valid primary key. If it is determined that the data record carries a primary key field, it matches the primary key field carried by the data record with the multiple pre-stored primary key fields locally. If it is determined that there is a primary key field in the multiple primary key fields that is the same as the primary key field carried by the data record, it is determined that the data record carries a valid primary key; otherwise, it is determined that the data record does not carry a valid primary key. Here, the multiple primary key fields pre-stored locally by the primary key generation system 200 can be the primary key fields of the corresponding data tables (or partitions) already established in the data integration system 300.

[0181] S2. The primary key generation system 200 determines whether it can generate a unified primary key for the data records without valid primary keys based on the initial broadcast variable. If it is determined that a unified primary key cannot be generated for the data records without valid primary keys, it stores the data records without valid primary keys in the streaming database and executes S3 - S6. If it is determined that a unified primary key can be generated for the data records without valid primary keys, it executes S7 - S9.

[0182] When initializing, the primary key generation system 200 generates an initial broadcast variable, which indicates that no high-confidence candidate keys have been mined (the high-confidence candidate keys are the Figure 4 first target candidate keys described in the embodiments).

[0183] Specifically, the primary key generation system 200 can determine whether the initial broadcast variable tells about high-confidence candidate keys. If it is determined that the initial broadcast variable does not tell about high-confidence candidate keys, it is determined that a unified primary key cannot be generated for the data records without valid primary keys; otherwise, it is determined that a unified primary key can be generated for the data records without valid primary keys.

[0184] S3. The primary key generation system 200 samples the data records in the streaming database.

[0185] S4. The primary key generation system 200 mines high-confidence candidate keys based on the sampled data records. If high-confidence candidate keys are mined, it executes S5. If no high-confidence candidate keys are mined, it executes S3 again to sample more data records in the streaming database and executes S4 based on the continuously sampled data records, and so on.

[0186] S5. The primary key generation system 200 writes the high-confidence candidate keys into the first broadcast variable.

[0187] S6. The primary key generation system 200 performs broadcast variable update, that is, uses the first broadcast variable to update the initial broadcast variable.

[0188] It can be understood that after the primary key generation system 200 performs broadcast variable update, in S2, it can determine whether a unified primary key can be generated for the newly received data record without a valid primary key based on the first broadcast variable.

[0189] S7. The primary key generation system 200 extracts the value of the corresponding high-confidence candidate key in the data record, and generates a unified primary key value for the data record based on the extracted value.

[0190] S8. The primary key generation system 200 adds the unified primary key (including the value of the unified primary key) to the data record and sends it to the data integration system 300 for integration.

[0191] In this example, when the primary key generation system 200 mines a high-confidence candidate key, as Figure 8 shown by the dashed arrow in, the primary key generation system 200 can also execute S7 - S8 for each data record in the streaming database, generate a unified primary key value for each data record, and then add the unified primary key (including the value of the unified primary key) to the data record and send it to the data integration system 300 for integration.

[0192] For the process of the above primary key generation system 200 mining high-confidence candidate keys based on sampled data records, reference can be made to Figure 5 the process of the primary key generation system described in S501 - S509 of the embodiment for determining target candidate keys (the first target candidate key, the second target candidate key,...) based on sampled data records. For the process of the primary key generation system 200 generating a unified primary key value for each data record, reference can be made to the relevant description in S504. For the sake of simplicity of the specification, it will not be elaborated here.

[0193] For the process of the data integration system 300 integrating multiple data records based on the unified primary key, reference can be made to the relevant description of S406 above.

[0194] It can be seen that through the above embodiments, the primary key generation system 200 can mine high-confidence candidate keys based on multiple data records without valid primary keys, and then generate a unified primary key for the multiple data records based on the mined high-confidence candidate keys, thus facilitating the data integration system 300 to integrate multiple data records based on the unified primary key.

[0195] For the convenience of more clearly understanding Figure 3 the specific process of generating a unified primary key for streaming data and integrating streaming data based on the unified primary key as shown in, the following combines Figure 9A schematic diagram of a streaming data integration process provided by an embodiment of the present application is introduced in more detail. It should be noted that in Figure 9 it is described by taking Figure 3 including two source ends as an example. As Figure 9 shown, the method includes the following steps:

[0196] S901: The client receives the first streaming data sent by the first source end.

[0197] S902: The client receives the second streaming data sent by the second source end.

[0198] S903: The client sends a primary key generation request to the primary key generation system. Correspondingly, the primary key generation system receives the primary key generation request sent by the client.

[0199] The primary key generation request carries the first streaming data and the second streaming data.

[0200] S904: The primary key generation system generates a unified primary key for the data record set.

[0201] S905: The primary key generation system adds the unified primary key generated for each data record to each data record.

[0202] S906: The primary key generation system sends each data record carrying the unified primary key to the data integration system. Correspondingly, the data integration system receives each data record carrying the unified primary key sent by the primary key generation system.

[0203] S907: The data integration system integrates each data record based on the unified primary key carried by each data record.

[0204] Figure 9 For concepts such as the first streaming data, the second streaming data, the data record set, and the unified primary key in Figure 4 please refer to the description of relevant concepts in the embodiment. For the sake of simplicity of the specification, no further elaboration is provided.

[0205] Combined with Figure 4 the steps shown, it can be seen that Figure 9 the steps shown are the same as or similar to Figure 4 the steps shown. Please refer to the specific description of the steps shown in Figure 4 For the sake of simplicity of the specification, no further elaboration is provided.

[0206] It should be noted that Figure 4 and Figure 9 the embodiment is described by taking generating a unified primary key for two streaming data as an example. The inventive concept provided by the present application can also be applied to generating a unified primary key for one or three or more streaming data. The specific implementation process is the same as Figure 4 andFigure 9 Similar to the embodiments, for the sake of brevity of the specification, it will not be elaborated further.

[0207] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0208] The above has elaborated in detail the primary key generation system and the primary key determination method provided by the embodiments of the present application. Based on the same inventive concept, the primary key determination device and the computing device cluster provided by the embodiments of the present application will be introduced below.

[0209] The primary key determination device can be applied to the above primary key generation system to implement various functions of the primary key generation system. The primary key determination device can be divided into one or more unit modules. Exemplarily, refer to Figure 10 , Figure 10 which is a schematic structural diagram of a primary key determination device provided by an embodiment of the present application. As shown in Figure 10 , it includes: an acquisition module 1010, a processing module 1030, and a sampling module 1020. The functions of each module of the primary key determination device 1000 will be introduced exemplarily below.

[0210] The acquisition module 1010 is configured to acquire a plurality of data records, and the plurality of data records include the same plurality of fields.

[0211] The sampling module 1020 is configured to sample the plurality of data records.

[0212] The processing module 1030 is configured to determine the fields and / or field combinations with duplicate values in the sampled data records as the first non-candidate keys.

[0213] The processing module 1030 is configured to determine the remaining elements in the universal set except the first non-candidate keys as a plurality of first possible candidate keys, where the universal set includes a plurality of fields and combinations of at least two fields in the plurality of fields.

[0214] The processing module 1030 is further configured to, when it is determined based on the plurality of data records that all the plurality of first possible candidate keys can be used as candidate keys, select a first target candidate key from the plurality of first possible candidate keys, or is further configured to, when it is determined based on the plurality of data records that at least one of the plurality of first possible candidate keys can be used as a candidate key, select a first target candidate key from the at least one first possible candidate key.

[0215] The processing module 1030 is further configured to generate a unified primary key value for each data record based on the value of the first target candidate key corresponding to each data record.

[0216] In some possible embodiments, none of the above-mentioned multiple data records carry a primary key, or, the multiple data records carry different primary keys, or, some of the multiple data records carry a primary key while the remaining ones do not.

[0217] In some possible embodiments, the above-mentioned sampling module 1020 is further configured to, when it is determined based on multiple data records that not all of the multiple first possible candidate keys can be used as candidate keys, continue to sample the multiple data records, and the above-mentioned processing module 1030 is further configured to perform the following operations:

[0218] Determine the fields and / or field combinations with duplicate values in the data records obtained by continuous sampling as the second non-candidate keys; determine the remaining elements in the universe except the first non-candidate keys and the second non-candidate keys as the multiple second possible candidate keys; when it is determined based on multiple data records that all of the multiple second possible candidate keys can be used as candidate keys, select a second target candidate key from the multiple second possible candidate keys; generate the value of the unified primary key for each data record based on the value of the corresponding second target candidate key in each data record.

[0219] In some possible embodiments, the above-mentioned processing module 1030 is configured to use the first possible candidate key with the most uniform data distribution among the multiple first possible candidate keys as the first target candidate key.

[0220] In some possible embodiments, the above-mentioned processing module 1030 is configured to hash the value of the corresponding first target candidate key in each data record, and use the obtained hash value as the value of the unified primary key for each data record.

[0221] In some possible embodiments, the above-mentioned processing module 1030 is configured to use the value of the corresponding first target candidate key in each data record as the value of the unified primary key for each data record.

[0222] In some possible embodiments, the above-mentioned acquisition module 1010 is further configured to acquire new data records, where the new data records also include multiple fields; the above-mentioned processing module 1030 is further configured to generate the value of the unified primary key for the new data records based on the value of the corresponding first target candidate key in the new data records.

[0223] In some possible embodiments, as Figure 10 shown, the apparatus 1000 further includes a sending module 1040; the above-mentioned processing module 1030 is further configured to add the value of the unified primary key of each data record to each data record; the above-mentioned sending module 1040 is configured to send each data record to the data integration system for integration.

[0224] In some possible embodiments, the above-mentioned multiple data records belong to the same data stream, or, belong to different data streams.

[0225] In some possible embodiments, when multiple data records belong to different data streams, the different data streams originate from the same source or different sources.

[0226] In some possible embodiments, after the obtaining module 1010 obtains multiple data records, the obtained multiple data records may be stored in the same data bucket (such as the first data bucket shown) Figure 10 Subsequently, the sampling module 1020 may sample the multiple data records from the first data bucket, and the processing module 1030 may determine multiple possible candidate keys (such as the multiple first possible candidate keys and the multiple second possible candidate keys) based on the sampled data records, and store the determined multiple possible candidate keys in another data bucket (such as the second data bucket shown) Figure 10 In addition, the processing module 1030 may also store the data records with the values of the unified primary key added in another data bucket (such as the third data bucket shown) Figure 10 Subsequently, the sending module 1040 may obtain the data records with the values of the unified primary key added from the third data bucket and send them to the data integration system. Optionally, the foregoing multiple data records, multiple possible candidate keys, and data records with the values of the unified primary key added may also all be stored in the same data bucket, or the multiple data records and multiple possible candidate keys are stored in the same data bucket, and the data records with the values of the unified primary key added are stored in another data bucket, or the multiple data records and the data records with the values of the unified primary key added are stored in the same data bucket, and the multiple possible candidate keys are stored in another data bucket. The present application does not make specific limitations thereon.

[0227] In specific implementation, the obtaining module 1010, the sampling module 1020, the processing module 1030, and the sending module 1040 may all be implemented by software or may be implemented by hardware. Exemplarily, next, taking the processing module 1030 as an example, the implementation manner of the processing module 1030 will be introduced. Similarly, the implementation manners of the obtaining module 1010, the sampling module 1020, and the sending module 1040 may refer to the implementation manner of the processing module 1030.

[0228] As an example of a software functional unit, the processing module 1030 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the processing module 1030 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.

[0229] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is achieved through the communication gateway.

[0230] As an example of a hardware functional unit, the processing module 1030 may include at least one computing device, such as a server. Alternatively, the processing module 1030 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0231] When the processing module 1030 includes multiple computing devices, the multiple computing devices included can be distributed in the same region or in different regions. The multiple computing devices included in the processing module 1030 can be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the processing module 1030 can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), and generic array logic (GALs).

[0232] It should be noted that in other embodiments, the processing module 1030 can be used to execute Figure 4 , Figure 5 , Figures 7 - 9 any of the steps executed by the primary key generation system in Figure 4 , Figure 5 , Figures 7 - 9 any of the steps executed by the primary key generation system in, the sampling module 1020 can be used to execute Figure 4 , Figure 5 , Figures 7 - 9 any of the steps executed by the primary key generation system in, the sending module 1040 can be used to execute Figure 4 , Figure 5 , Figures 7 - 9 any of the steps executed by the primary key generation system in. The steps to be implemented by the obtaining module 1010, the processing module 1030, the sampling module 1020, and the sending module 1040 can be specified as needed. By the obtaining module 1010, the processing module 1030, the sampling module 1020, and the sending module 1040, different steps in Figure 4 , Figure 5 , Figures 7 - 9 executed by the primary key generation system are respectively implemented to realize all functions of the primary key determination device 1000.

[0233] It should be understood that the functions of the above-described respective modules are only the functions that the primary key determination device 1000 can have in some embodiments of the present application. The present application does not limit the functions of the respective modules.

[0234] It should also be understood that Figure 10 is an exemplary partitioning method. The primary key determination device 1000 can also include more or fewer modules. Specifically, the partitioning method of the modules in the primary key determination device 1000 can be flexibly adjusted based on the actual business scenario, and the present application does not make specific limitations.

[0235] An embodiment of the present application further provides a computing device 1100, and this computing device 1100 can be deployedFigure 10 The primary key determination device 1000 shown, which can also be the aforementioned primary key generation system, and the operations and / or functions of the various modules in the computing device 1100 are respectively for implementing Figure 4 、 Figure 5 、 Figures 7 - 9 the corresponding steps executed by the primary key generation system in the method shown.

[0236] As Figure 11 shown, the computing device 1100 includes: a processor 1110, a memory 1120, and a communication interface 1130. Among them, the processor 1110, the memory 1120, and the communication interface 1130 can be interconnected with each other through a bus 1140.

[0237] The processor 1110 can read the program code (including instructions) stored in the memory 1120 and execute the program code stored in the memory 1120, so that the computing device 1100 executes Figure 4 、 Figure 5 、 Figures 7 - 9 the steps executed by the primary key generation system in , or cause the computing device 1100 to deploy the primary key determination device 1000.

[0238] The processor 1110 can have various specific implementation forms. For example, it can be a central processing unit (CPU), or a combination of a CPU and a hardware chip. The above hardware chip can be an ASIC, a PLD, or a combination thereof. The above PLD can be a CPLD, an FPGA, a GAL, or any combination thereof. The processor 1110 executes various types of digital storage instructions, such as software or firmware programs stored in the memory 1120, and it can enable the computing device 1100 to provide a wide variety of services.

[0239] In a specific implementation, as an embodiment, the processor 1110 includes one or more CPUs.

[0240] In a specific implementation, as an embodiment, the computing device 1100 also includes multiple processors, and each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor refers to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0241] The memory 1120 is used to store program code and is controlled by the processor 1110 to execute the above Figure 4 、 Figure 5 、 Figures 7 - 9 the steps executed by the primary key generation system in. The program code can include one or more software modules, and these one or more software modules can beFigure 10 The software modules provided in the embodiments, such as the acquisition module 1010, the processing module 1030, the sampling module 1020, and the sending module 1040.

[0242] The memory 1120 may include volatile memory, such as random access memory (RAM); the memory 1120 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); the memory 1120 may further include a combination of the above types.

[0243] The communication interface 1130 may be a wired interface (such as an Ethernet interface, a fiber optic interface, other types of interfaces (e.g., InfiniBand interface)) or a wireless interface (such as a cellular network interface or a wireless local area network interface) for communicating with other computing devices or apparatuses. The communication interface 1130 may adopt a protocol family over Transmission Control Protocol / Internet Protocol (TCP / IP), such as, for example, Remote Function Call (RFC) protocol, Simple Object Access Protocol (SOAP) protocol, Simple Network Management Protocol (SNMP) protocol, Common Object Request Broker Architecture (CORBA) protocol, and distributed protocols, etc.

[0244] The bus 1140 can be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The bus 1140 can be divided into an address bus, a data bus, a control bus, etc.

[0245] In addition to the data bus, the bus 1140 can also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clear illustration, all kinds of buses are labeled as the bus 1140 in the figure. For the sake of easy representation, Figure 11 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0246] The above computing device 1100 is used to execute Figure 4 、 Figure 5 、 Figures 7 - 9 the steps executed by the primary key generation system in the above, and the specific implementation process can be seen in the above method embodiments, which will not be elaborated here.

[0247] It should be understood that the computing device 1100 is only an example provided in the embodiments of the present application, and, the computing device 1100 may have more or fewer components than Figure 11 the components shown, two or more components can be combined, or different configurations of components can be implemented. For the content not shown or described in the embodiments of the present application, reference can be made to the relevant descriptions in the foregoing Figures 1 - 10 embodiments, which will not be elaborated here.

[0248] The present application also provides a computing device cluster 1200, and the computing device cluster 1200 can deploy Figure 10 the primary key determination device 1000 shown, or can be the foregoing primary key generation system 200. The operations and / or functions of each module in the computing device cluster 1200 are respectively to implement Figure 4 、 Figure 5 、 Figures 7 - 9 the corresponding steps executed by the primary key generation system in the method shown.

[0249] Such as Figure 12As shown, the computing device cluster 1200 includes at least one computing device 1100. In the memory 1120 of one or more computing devices 1100 in the computing device cluster, there may be stored the same instructions for executing Figure 4 , Figure 5 , Figures 7 - 9 the steps executed by the primary key generation system. The computing device 1100 may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device 1100 may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0250] In some possible implementation manners, in the memory 1120 of one or more computing devices 1100 in the computing device cluster 1200, there may also be respectively stored instructions for executing Figure 4 , Figure 5 , Figures 7 - 9 the steps executed by the primary key generation system. In other words, a combination of one or more computing devices 1100 may jointly execute the instructions for executing Figure 4 , Figure 5 , Figures 7 - 9 the steps executed by the primary key generation system.

[0251] It should be noted that the memories 1120 in different computing devices 1100 in the computing device cluster 1200 may store different instructions, respectively for executing partial functions of the primary key determination device 1000. That is, the instructions stored in the memories 1120 of different computing devices 1100 may implement the functions of one or more of the acquisition module 1010, the processing module 1030, the sampling module 1020, and the sending module 1040.

[0252] In some possible implementation manners, one or more computing devices 1100 in the computing device cluster 1200 may be connected through a network. Among them, the network may be a wide area network or a local area network, etc. Figure 13 A possible implementation manner is shown. As Figure 13 shown, two computing devices 1100A and 1100B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation manners, in the memory 1120 of the computing device 1100A, there are stored instructions for executing the functions of the acquisition module 1010 and the sending module 1040. At the same time, in the memory 1120 of the computing device 1100B, there are stored instructions for executing the functions of the processing module 1030 and the sampling module 1020.

[0253] Figure 13The connection mode between the computing device clusters 1200 shown can be considered as follows. Since the primary key determination method provided in the embodiments of the present application needs to sample a large number of data records, analyze the sampled data records, and generate a unified primary key for a large number of data records, it is considered to execute the functions implemented by the sampling module 1020 and the processing module 1030 on the computing device 1100B.

[0254] It should be understood that Figure 13 the functions of the computing device 1100A shown in can also be completed by multiple computing devices 1100. Similarly, the functions of the computing device 1100B can also be completed by multiple computing devices 1100.

[0255] The present application also provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute Figure 4 、 Figure 5 、 Figures 7 - 9 the steps executed by the primary key generation system in.

[0256] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a high-density digital video disc (DVD)), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute Figure 4 、 Figure 5 、 Figures 7 - 9 the steps executed by the primary key generation system in.

[0257] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0258] In the above embodiments, it can be implemented in whole or in part by software, hardware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium, etc.

[0259] As described above, the foregoing is only a specific implementation manner of the present application. Those skilled in the art, based on the specific implementation manner provided in the present application, can think of variations or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A primary key determination method, characterized in that: The method comprises: The primary key generation system obtains a plurality of data records, wherein the plurality of data records include the same plurality of fields; Sampling the multiple data records, and determining fields and / or field combinations with duplicate values ​​in the sampled data records as first non-candidate keys; Determine the remaining elements in the full set except the first non-candidate key as a plurality of first possible candidate keys, wherein the full set includes the plurality of fields and a combination of at least two of the plurality of fields; If it is determined based on the multiple data records that all of the multiple first possible candidate keys can be used as candidate keys, a first target candidate key is selected from the multiple first possible candidate keys; or, if it is determined based on the multiple data records that at least one of the multiple first possible candidate keys can be used as a candidate key, a first target candidate key is selected from the at least one first possible candidate key; Based on the value of the first target candidate key in each data record, a value of the unified primary key of each data record is generated.

2. The method according to claim 1, characterized in that None of the multiple data records carry a primary key, or the multiple data records carry different primary keys, or some of the multiple data records carry a primary key and the rest do not carry a primary key.

3. The method according to claim 1 or 2, characterized in that: The method further comprises: If it is determined based on the multiple data records that the multiple first possible candidate keys cannot all be candidate keys, continue sampling the multiple data records, and determine fields and / or field combinations with repeated values ​​in the data records to be sampled as second non-candidate keys; Determine the remaining elements in the full set except the first non-candidate key and the second non-candidate key as a plurality of second possible candidate keys; When it is determined based on the multiple data records that the multiple second possible candidate keys can all be used as the candidate key, selecting a second target candidate key from the multiple second possible candidate keys; Based on the value of the second target candidate key corresponding to each data record, the value of the second target candidate key of the unified primary key of each data record is generated.

4. The method according to any one of claims 1 to 3, characterized in that: The step of selecting a first target candidate key from the plurality of first possible candidate keys comprises: The first possible candidate key with the most uniform data distribution among the multiple first possible candidate keys is used as the first target candidate key.

5. The method according to any one of claims 1 to 4, characterized in that: The generating, based on the value of the first target candidate key in each data record, the value of the unified primary key of each data record includes: A hash is performed on the value corresponding to the first target candidate key in each data record, and the obtained hash value is used as the value of the unified primary key of each data record.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: When a new data record is received, a value of a unified primary key of the new data record is generated based on a value corresponding to the first target candidate key in the new data record, wherein the new data record also includes the multiple fields.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Adding the value of the unified primary key of each data record to each data record; Each of the data records is sent to a data integration system for integration.

8. The method according to any one of claims 1 to 7, characterized in that: The multiple data records belong to the same data stream, or belong to different data streams.

9. The method according to claim 8, characterized in that In the case that the multiple data records belong to different data streams, the different data streams originate from the same source end or different source ends.

10. A primary key determination device, characterized in that: The device comprises: An acquisition module, used for acquiring a plurality of data records, wherein the plurality of data records include a plurality of identical fields; A sampling module, configured to sample the plurality of data records; A processing module, configured to determine fields and / or field combinations having duplicate values ​​in the sampled data records as first non-candidate keys; The processing module is used to determine the remaining elements in the full set except the first non-candidate key as a plurality of first possible candidate keys, wherein the full set includes the plurality of fields and a combination of at least two of the plurality of fields; The processing module is further configured to select a first target candidate key from the plurality of first possible candidate keys when it is determined based on the plurality of data records that all of the plurality of first possible candidate keys can be used as candidate keys, or select a first target candidate key from the at least one first possible candidate key when it is determined based on the plurality of data records that at least one of the plurality of first possible candidate keys can be used as a candidate key; The processing module is further configured to generate a value of a unified primary key for each data record based on a value corresponding to the first target candidate key in each data record.

11. The device according to claim 10, characterized in that None of the multiple data records carry a primary key, or the multiple data records carry different primary keys, or some of the multiple data records carry a primary key and the rest do not carry a primary key.

12. The device according to claim 10 or 11, characterized in that The sampling module is further configured to continue sampling the plurality of data records if it is determined based on the plurality of data records that the plurality of first possible candidate keys cannot all be candidate keys; The processing module is further used to determine the fields and / or field combinations with repeated values ​​in the continuously sampled data records as the second non-candidate key; The processing module is further used to determine the remaining elements in the full set except the first non-candidate key and the second non-candidate key as a plurality of second possible candidate keys; The processing module is further configured to select a second target candidate key from the plurality of second possible candidate keys when it is determined based on the plurality of data records that the plurality of second possible candidate keys can all serve as the candidate key; The processing module is further configured to generate a value of a unified primary key for each data record based on a value corresponding to the second target candidate key in each data record.

13. The device according to any one of claims 10 to 12, characterized in that The processing module is used to use the first possible candidate key with the most uniform data distribution among the multiple first possible candidate keys as the first target candidate key.

14. The device according to any one of claims 10 to 13, characterized in that The processing module is used to hash the value corresponding to the first target candidate key in each data record, and use the obtained hash value as the value of the unified primary key of each data record.

15. The device according to any one of claims 10 to 14, characterized in that The acquisition module is further used to acquire a new data record, wherein the new data record also includes the multiple fields; The processing module is further configured to generate a value of a unified primary key of the new data record based on a value corresponding to the first target candidate key in the new data record.

16. The device according to any one of claims 10 to 15, characterized in that The device also includes: a sending module; The processing module is further used to add the value of the unified primary key of each data record to each data record; The sending module is used to send each data record to the data integration system for integration.

17. The device according to any one of claims 10 to 16, characterized in that The multiple data records belong to the same data stream, or belong to different data streams.

18. The device according to claim 17, characterized in that In the case that the multiple data records belong to different data streams, the different data streams originate from the same source end or different source ends.

19. A computing device cluster, characterized in that: It includes at least one computing device, each of which includes a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 9.