Primary-key determination method and apparatus, and computer-readable storage medium
By determining candidate keys in streaming data and generating unified primary keys, the complexity problem of streaming data integration is solved, and efficient data integration and query is achieved.
Patent Information
- Application Number
- PCT/CN2024/108017
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-22
- Filing Date
- 2024-07-27
- Publication Date
- 2025-05-22
AI Technical Summary
The primary key of streaming data is complex, resulting in the inability to integrate.
A primary key determination method is provided, by obtaining multiple data records, sampling and determining candidate keys, selecting appropriate candidate keys to generate a unified primary key for each data record.
It realizes the integration of streaming data, improves the query performance and indexing efficiency of data, and ensures the security and maintainability of data.
Smart Images

Figure CN2024108017_22052025_PF_FP_ABST
Abstract
Description
Primary key determination method, device, and computer-readable storage medium
[0001] This application claims priority to the Chinese patent application with application number 202311541738.4 filed with the State Intellectual Property Office of China on November 17, 2023, and priority to the Chinese patent application with invention name “Streaming data processing method, device and computer-readable storage medium”, as well as priority to the Chinese patent application with application number 202410088939.1 filed with the State Intellectual Property Office of China on January 22, 2024, and priority to the Chinese patent application with invention name “Primary key determination method, device and computer-readable storage medium”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of data processing technology, and in particular to a primary key determination method, device, and computer-readable storage medium. Background Art
[0003] Streaming data refers to data generated in a continuous time series. This type of data is highly timely, reflects changes in business status, and contains a large amount of real-time information and event records. For example, meteorological observation data records collected in real time by meteorological observation stations, employee clock-in records collected in real time by enterprise clock-in systems, user login records collected in real time by enterprise service systems, and commodity transaction records collected in real time by enterprise transaction systems. Real-time integrated processing of streaming data can effectively assist in business decision-making.
[0004] However, the primary key situation of streaming data is very complicated. For example, streaming data may not have a primary key, or some data may have a primary key while others may not, or the primary keys may be different, making integration impossible.
[0005] Summary of the Invention
[0006] The present application provides a primary key determination method, device, and computer-readable storage medium, which can implement streaming data integration with very complex primary key situations.
[0007] In a first aspect, a method for determining a primary key is provided, the method comprising the following steps:
[0008] A primary key generation system obtains multiple data records, each of which includes the same multiple fields; samples the multiple data records, and determines the fields and / or field combinations with duplicate values in the sampled data records as first non-candidate keys; determines the remaining elements in the full set except the first non-candidate key as multiple first possible candidate keys, wherein the full set includes multiple fields and a combination of at least two of the multiple fields; when it is determined based on the multiple data records that multiple first possible candidate keys can all be used as candidate keys, a first target candidate key is selected from the multiple first possible candidate keys, or when it is determined based on the multiple data records that at least one first possible candidate key among the multiple first possible candidate keys can be used as a candidate key, a first target candidate key is selected from at least one first possible candidate key; and generates a unified primary key value for each data record based on the value of the corresponding first target candidate key in each data record.
[0009] The above-mentioned multiple data records belong to the same data stream (also called streaming data), or belong to different data streams.
[0010] In some possible implementations, the method further includes the following steps:
[0011] The primary key generation system adds the value of the unified primary key of each data record to each data record and sends each data record to the data integration system for integration.
[0012] The above scheme can analyze multiple data records that include the same multiple fields in the streaming data, determine possible candidate keys, and select the target candidate key after verifying the possible candidate keys. Then, based on the values of the corresponding target candidate keys in the multiple data records, a unified primary key value is generated for the multiple data records. In this way, the subsequent data integration system can integrate multiple data records based on the unified primary key of multiple data records, thereby realizing the integration of streaming data.
[0013] In some possible implementations, none of the multiple data records carry a primary key, or the multiple data records carry different primary keys, or some of the multiple data records carry a primary key and the rest do not.
[0014] In some possible implementations, the above method further includes the following steps:
[0015] When the primary key generation system determines based on multiple data records that multiple first possible candidate keys cannot all be candidate keys, the system continues to sample multiple data records, and determines the fields and / or field combinations with duplicate values in the data records to be sampled as second non-candidate keys; determines the remaining elements in the full set except the first non-candidate key and the second non-candidate key as multiple second possible candidate keys; when it is determined based on multiple data records that multiple second possible candidate keys can all be candidate keys, the system selects a second target candidate key from the multiple second possible candidate keys; and generates a unified primary key value for each data record based on the value corresponding to the second target candidate key in each data record.
[0016] By analogy based on the above scheme, when the primary key generation system determines that multiple second possible candidate keys cannot all be candidate keys based on multiple data records, it can continue to sample multiple data records, and determine the fields and / or field combinations with repeated values in the data records to be sampled as third non-candidate keys,..., and generate a unified primary key value for each data record.
[0017] It can be seen that the above solution can filter out all non-candidate keys from the entire set, leaving all candidate keys, and select appropriate candidate keys from the remaining candidate keys, and generate a unified primary key value for each data record based on the appropriate candidate keys.
[0018] In some possible implementations, the primary key generation system may select the first target candidate key from multiple first possible candidate keys in the following manner: select the first possible candidate key with the most even data distribution among the multiple first possible candidate keys as the first target candidate key.
[0019] In a database, primary keys are usually used for indexing to speed up data retrieval. If the primary key values are evenly distributed, the indexing effect will be better and the query speed will be faster.
[0020] In the above scheme, since the data distribution of the first target candidate key selected from multiple first possible candidate keys is uniform, it can be understood that this can make the distribution of the value of the unified primary key generated for multiple data records based on the value of the first target candidate key also relatively uniform. Then, when the data integration system uses the value of the unified primary key as the index of multiple data records, the indexing effect will be better and the query speed will be faster.
[0021] In some possible implementations, the primary key generation system may hash the value corresponding to the first target candidate key in each data record, and use the obtained hash value as the value of the unified primary key of each data record.
[0022] Optionally, the primary key generation system may also use the value corresponding to the first target candidate key in each data record as the value of the unified primary key of each data record.
[0023] It can be understood that the primary key generation system hashes the value corresponding to the first target candidate key in each data record and uses the obtained hash value as the value of the unified primary key of each data record. Compared with using the value corresponding to the first target candidate key in each data record as the value of the unified primary key of each data record, since the hash function is a one-way function, that is, the original value cannot be deduced from the hash value, it can avoid direct exposure of the value corresponding to the first target candidate key in each data record, thereby improving data security.
[0024] Furthermore, because hash functions typically map input data uniformly to the output space, this ensures that the generated hash values are relatively evenly distributed within the unified primary key. Uniformly distributed primary keys help improve data query performance and indexing efficiency. Furthermore, because hash values generated by hash functions typically have a fixed length and are unaffected by the length of the input data, this ensures that the unified primary key value occupies a fixed amount of space during storage and indexing, improving database storage efficiency and query performance. Furthermore, because the hash value is calculated based on the value of the first target candidate key in each data record and is independent of the value of the first target candidate key in each data record, the generated hash value remains unchanged even if the value of the first target candidate key in each subsequent data record changes. Therefore, the value of the primary key field remains unchanged. This is very useful for data updates and maintenance, as it avoids primary key conflicts and data consistency issues caused by changes to the original data. In short, using hash values as unified primary key values offers advantages over directly using the original data as unified primary key values, including data protection, uniqueness assurance, uniform distribution, fixed length, and irrelevance. These advantages can improve database security, performance, and maintainability.
[0025] In some possible implementations, the above method further includes the following steps:
[0026] When receiving a new data record, the primary key generation system generates a unified primary key value for the new data record based on the value corresponding to the first target candidate key in the new data record, wherein the new data record also includes multiple fields.
[0027] In some possible implementations, when multiple data records belong to different data streams, the different data streams originate from the same source end or different source ends.
[0028] It can be seen that the primary key determination method provided in this application can generate a unified primary key for the same-source streaming data, and can also generate a unified primary key for cross-source streaming data.
[0029] In a second aspect, a primary key determination device is provided, the device comprising:
[0030] An acquisition module, configured to acquire a plurality of data records, wherein the plurality of data records include the same plurality of fields;
[0031] a sampling module, configured to sample the plurality of data records;
[0032] A processing module, configured to determine fields and / or field combinations with duplicate values in the sampled data records as first non-candidate keys;
[0033] The processing module is configured to determine the remaining elements in the full set except the first non-candidate key as a plurality of first possible candidate keys, wherein the full set includes the plurality of fields and a combination of at least two of the plurality of fields;
[0034] The processing module is further configured to, if it is determined based on the multiple data records that all of the multiple first possible candidate keys can be used as candidate keys, select a first target candidate key from the multiple first possible candidate keys, or, if it is determined based on the multiple data records that at least one of the multiple first possible candidate keys can be used as a candidate key, select a first target candidate key from the at least one first possible candidate key;
[0035] The processing module is further configured to generate a value of a unified primary key for each data record based on the data corresponding to the first target candidate key in each data record.
[0036] In some possible implementations, none of the multiple data records carry a primary key, or the multiple data records carry different primary keys, or some of the multiple data records carry a primary key and the rest do not.
[0037] In some possible implementations, the sampling module is further configured to, if it is determined based on the multiple data records that all of the multiple first possible candidate keys cannot be used as candidate keys, continue sampling the multiple data records;
[0038] The processing module is further configured to determine fields and / or field combinations with repeated values in the continuously sampled data records as second non-candidate keys;
[0039] The processing module is further configured to determine the remaining elements in the full set except the first non-candidate key and the second non-candidate key as a plurality of second possible candidate keys;
[0040] The processing module is further configured to select a second target candidate key from the plurality of second possible candidate keys if it is determined based on the plurality of data records that the plurality of second possible candidate keys can all serve as the candidate key;
[0041] The processing module is further configured to generate a unified primary key value for each data record based on a value corresponding to the second target candidate key in each data record.
[0042] In some possible implementations, the processing module is configured to use a first possible candidate key with the most uniform data distribution among the multiple first possible candidate keys as the first target candidate key.
[0043] In some possible implementations, the processing module is configured to hash the data corresponding to the first target candidate key in each data record, and use the obtained hash value as the value of the unified primary key of each data record.
[0044] In some possible implementations, the acquisition module is also used to acquire a new data record, wherein the new data record also includes the multiple fields; the processing module is also used to generate a unified primary key value for the new data record based on the value corresponding to the first target candidate key in the new data record.
[0045] In some possible implementations, the device also includes: a sending module; the processing module is also used to add the value of the unified primary key of each data record to each data record; the sending module is used to send each data record to the data integration system for integration.
[0046] In some possible implementations, the multiple data records belong to the same data stream, or belong to different data streams.
[0047] In some possible implementations, when the multiple data records belong to different data streams, the different data streams originate from the same source end or different source ends.
[0048] In a third aspect, a computing device is provided, comprising a processor and a memory, wherein the memory is used to store instructions and the processor is used to execute instructions, so that the computing device implements the method described in the first aspect and any implementation manner of the first aspect.
[0049] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions, and when the instructions are executed by a computing device or a computing device cluster, the method described in the first aspect and any implementation manner of the first aspect is implemented.
[0050] In a fifth aspect, a computing device cluster is provided, which includes at least one computing device, each of the at least one computing device includes a processor and a memory, and the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster implements the method described in the first aspect and any implementation method of the first aspect.
[0051] In a sixth aspect, a computer program product comprising instructions is provided, wherein the computer program product includes instructions that can be run on a computing device or stored in software or program products in any available medium. When the computer program product is run on a computing device or a computing device cluster, the computing device or computing device cluster executes the method described in the first aspect and any implementation method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] FIG1 is a schematic diagram of the architecture of a meteorological observation data integration system according to an embodiment of the present application;
[0053] FIG2 is a schematic diagram of the architecture of a primary key generation system provided in an embodiment of the present application;
[0054] FIG3 is a schematic diagram of the architecture of another primary key generation system provided in an embodiment of the present application;
[0055] FIG4 is a schematic diagram of a streaming data integration process provided in an embodiment of the present application;
[0056] FIG5 is a schematic diagram of a process of generating a unified primary key for a data record set by a primary key generation system provided in an embodiment of the present application;
[0057] FIG6 is a schematic diagram of selecting a first target candidate key with uniform data distribution through a histogram according to an embodiment of the present application;
[0058] 7 is a schematic diagram of a process for generating a unified primary key for a data record set by another primary key generation system provided in an embodiment of the present application;
[0059] FIG8 is a schematic diagram of another streaming data integration process provided by an embodiment of the present application;
[0060] FIG9 is a schematic diagram of a streaming data integration process provided in an embodiment of the present application;
[0061] FIG10 is a schematic structural diagram of a primary key determination device provided in an embodiment of the present application;
[0062] FIG11 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0063] FIG12 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;
[0064] FIG13 is a schematic diagram of the structure of another computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION
[0065] First, an application scenario involved in an embodiment of the present application is introduced.
[0066] Taking the application scenario of meteorological observation data integration as an example, as shown in FIG1 , the meteorological observation data integration scenario includes multiple meteorological observation stations 10 and a data integration system 20 .
[0067] A plurality of meteorological observation stations 10 are located in different areas and are used to periodically collect meteorological observation data such as atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, rainfall and ultraviolet index in the area where they are located, and transmit the meteorological observation data to the data integration system 20 in real time as streaming data. If it is necessary to perform persistent storage or association operations on the streaming data, a unique primary key value can be assigned to each data record of the streaming data to facilitate subsequent processing and analysis of the streaming data. The primary key can be data unrelated to the streaming data, such as the time when the streaming data is collected (i.e., the collection time), or it can be data related to the streaming data, such as atmospheric temperature, etc. Since the streaming data is generated by a plurality of different meteorological observation stations 10, the units to which the different meteorological observation stations belong, the collection equipment used, the communication method, etc. may not be the same, and therefore, the situation of the primary key of the streaming data is very complicated. The situation of the primary key of the streaming data can be divided into the following situations:
[0068] Case 1: No primary key exists in the streaming data.
[0069] Case ②: All streaming data have primary keys, but since the streaming data come from different source ends (i.e., different meteorological observation stations 10), their primary keys may be different. For example, the primary key carried by the streaming data of meteorological observation station 101 is the time when the streaming data is collected, while the primary keys carried by the streaming data of meteorological observation stations 102-108 are the time when the streaming data is collected and the collected atmospheric temperature.
[0070] Case ③: Part of the streaming data does not have a primary key, while the remaining part has a primary key. For example, the streaming data of meteorological observation station 102 does not have a primary key, while the streaming data of meteorological observation stations 102-108 all have primary keys.
[0071] The data integration system 20 is used to integrate streaming data received in real time from multiple meteorological observation stations 10. The data integration system 20 can be a data lake, a data middle platform or a real-time data warehouse, etc. Integration refers to integrating streaming data from a certain data source or from different data sources but belonging to the same business type into a unified data storage (such as a unified data table or a unified partition) for analysis, query and application development operations. Integration aims to solve problems such as dispersion, redundancy and inconsistency of streaming data, so that streaming data can be more conveniently accessed and utilized. Specifically, the data integration system 20 can perform integration based on the primary key of the streaming data, and integrate streaming data with the same primary key into a unified data storage.
[0072] It should be understood that FIG1 is merely an example of a meteorological observation data integration system. The number of meteorological observation stations 10 and data integration systems 20 can be one or any number of meteorological observation stations 10 and data integration systems 20, and this application does not impose any specific limitation.
[0073] It can be understood that in addition to the above-mentioned meteorological observation data integration scenario, there are other application scenarios, such as the enterprise employee punch-in record integration scenario. The streaming data in this scenario includes employee ID, employee name, punch-in time, punch-in location, punch-in method, punch-in type, overtime information, abnormal conditions, etc., among which the punch-in method can be swiping card, fingerprint recognition, facial recognition, mobile phone application, etc., the punch-in type can be punching in at work, punching out at get off work, punching in for overtime, etc., the overtime information can include whether overtime is worked and the start and end time of overtime, and abnormal conditions can be lateness, early leaving, etc.; for example, the user login record integration scenario of various service systems of an enterprise, the streaming data in this scenario includes user name, password, name, role / authority, email address, mobile phone number, department / organization, login time, login Internet Protocol address (internet protocol address, referred to as IP address), etc.; for another example, in the commodity transaction record integration scenario of various transaction systems of an enterprise, the streaming data in this scenario includes commodity transaction records, which may include the purchase account name, purchased commodity type, purchased commodity quantity, paid amount, order number, order creation time, order payment time, order shipment time, and order receipt time, etc.
[0074] However, due to the complexity of the primary key of the streaming data, the data integration system 20 cannot perform integration.
[0075] In response to the above problems, the present application provides a primary key generation system, a primary key determination method and device, etc., which can generate a unified primary key for multiple streaming data by analyzing multiple streaming data when multiple streaming data do not have a primary key, or when some of the multiple streaming data have a primary key and some do not, or when the primary keys of multiple streaming data are different, so that the data integration system can integrate multiple streaming data based on the unified primary key.
[0076] Next, the primary key generation system, primary key determination method and device, etc. provided by the application will be described in detail with reference to the corresponding drawings.
[0077] Please refer to Figure 2, which is a schematic diagram of the architecture of a primary key generation system provided in an embodiment of the present application. As shown in Figure 2, the architecture includes a client 100, a primary key generation system 200 and a data integration system 300.
[0078] The client 100 can be deployed at the source end (the meteorological observation station 10 in the meteorological observation data integration scenario as shown in Figure 1, the employee punch-in record collection device (such as a punch-in machine) in the enterprise employee punch-in record integration scenario, the user login record collection device in the user login record integration scenario of each enterprise service system, the commodity transaction record collection device in the commodity transaction record integration scenario of each enterprise transaction system, etc.).
[0079] In a specific implementation, the client 100 can be used to realize human-computer interaction. It can be a software or application running on a user-controlled device (such as the above-mentioned meteorological observation station 10, employee clock-in record collection device, user login record collection device, commodity transaction record collection device) or a computing device, such as a personal computer client, or a browser-based World Wide Web client, or an application (APP) client running on a mobile terminal. This application does not make specific limitations.
[0080] Alternatively, the client 100 may be a separate client specifically used to implement streaming data primary key generation, such as a primary key generation tool or application. Alternatively, the client 100 may be a primary key generation module or plug-in within comprehensive software, such as a streaming data primary key generation module within commonly used enterprise data integration software, and this application does not impose any specific limitations.
[0081] Optionally, the client 100 may also be a client of a cloud platform, such as a console of the cloud platform, which may be a web-based console or an application programming interface (API)-based console, which is not specifically limited in this application. The console may provide users with a streaming data primary key generation cloud service, and users may obtain access to the primary key generation system 200 provided in this application by purchasing the cloud service.
[0082] The primary key generation system 200 can be deployed on a computing device or a cluster of computing devices. Computing devices include bare metal servers (BMS), virtual machines (VM), containers, or edge computing devices. BMS refers to a general-purpose physical server, such as an ARM server or an X86 server; a virtual machine refers to a complete computer system with complete hardware system functions that is simulated by software and runs in a completely isolated environment. Any work that can be done in a physical computer can be done in a virtual machine. When creating a virtual machine in a computing device, part of the physical machine's hard disk and memory capacity needs to be used as the virtual machine's hard disk and memory capacity. Each virtual machine has an independent basic input / output system (BIOS), hard disk, and operating system, and can be operated like a physical machine. A container is a portable software unit that can combine an application and all its dependencies into a software package that is not restricted by the underlying host operating system. This eliminates the need to build a complex environment and simplifies the process from application development to deployment. Edge computing devices refer to devices that are closer to data sources and end users and have low latency and high bandwidth characteristics, such as smart routers, edge servers, etc. The computing device cluster may include multiple computing devices as described above, such as a data center, which is not specifically limited in this application.
[0083] The data integration system 300 can be deployed on a computing device or a computing device cluster. The introduction of the computing device or computing device cluster is as mentioned above.
[0084] Optionally, the primary key generation system 200 can be deployed in the same computing device as the data integration system 300, or deployed in different computing devices in the same computing device cluster, or deployed in computing devices and storage arrays in the same computing device cluster, or deployed in different computing device clusters. This application does not make specific limitations.
[0085] Optionally, the primary key generation system 200 can be deployed in the same or different computing device clusters as the client 100, for example, the client 100 is deployed on a computing device in a first computing device cluster, and the primary key generation system 200 is deployed on a computing device in a second computing device cluster; or, the client 100 and the primary key generation system 200 are deployed in the same computing device cluster. It should be understood that the above examples are for illustration only and are not specifically limited in this application.
[0086] In other possible implementations, all functions of the above-mentioned primary key generation system 200 can also be implemented by the source end that generates the streaming data. For example, the source end implements the above-mentioned primary key generation service to generate a unified primary key for the streaming data generated by itself, or the source end implements the above-mentioned primary key generation service to generate a unified primary key for the streaming data generated by other source ends.
[0087] There is a communication connection between the client 100, the primary key generation system 200 and the data integration system 300, which can be a wired connection or a wireless connection, and this application does not make any specific limitations.
[0088] It should be understood that the architecture shown in Figure 2 is only for example. For example, in actual applications, including network equipment for forwarding communication data between the primary key generation system 200, the client 100 and the data integration system 300, the number of clients 100 / data integration systems 300 that establish communication connections with the primary key generation system 200 can be one or more, and this application does not make specific restrictions. The number of primary key generation systems 200 can also be one or more, and this application does not make specific restrictions.
[0089] Please refer to Figure 3, which is a schematic diagram of the architecture of another primary key generation system provided in an embodiment of the present application. As shown in Figure 3, the architecture includes a client 100, a primary key generation system 200 and a data integration system 300.
[0090] 2, the architecture shown in FIG3 differs from the architecture shown in FIG2 in that the client 100 is deployed independently of the source end. Specifically, the client 100 can be deployed on a computing device or a computing device cluster, and the introduction of the computing device or computing device cluster is referred to above.
[0091] Regarding the similarities between the architecture shown in FIG3 and the architecture shown in FIG2 , please refer to the relevant description in FIG2 and no further details will be given.
[0092] It should be understood that the architecture shown in Figure 3 is only for example. For example, in actual applications, including network equipment for forwarding communication data between the primary key generation system 200, the client 100 and the data integration system 300, the number of clients 100 / data integration systems 300 that establish communication connections with the primary key generation system 200 can be one or more, and this application does not make specific restrictions. The number of primary key generation systems 200 can also be one or more, and this application does not make specific restrictions.
[0093] In order to more clearly understand the specific process of generating a unified primary key for multiple stream data and integrating multiple stream data based on the unified primary key in the architecture shown in Figure 2, a more detailed introduction is given below in conjunction with a schematic diagram of a stream data integration process provided by an embodiment of the present application shown in Figure 4. It should be noted that in Figure 4, the example of Figure 2 including two source terminals and each source terminal deploying a client is used for illustration. As shown in Figure 4, the method includes the following steps:
[0094] S401: A first client obtains first streaming data.
[0095] The first streaming data is a data stream acquired by the first client in a time period. The first streaming data may be meteorological observation data, employee clock-in records, user login records, commodity transaction records, etc. This application does not limit the first streaming data. The first stream data includes multiple fields, and when the first stream data is meteorological observation data, the fields of the first stream data include one or more of atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, and rainfall, etc. When the first stream data is an employee punch-in record, the fields of the first stream data include one or more of employee work number, employee name, punch-in time, punch-in location, punch-in method, punch-in type, overtime information, abnormal situation, etc. When the first stream data is an employee punch-in record, the fields of the first stream data include one or more of purchase account name, purchase product type, purchase product quantity, paid amount, order number, order creation time, order payment time, order shipment time, and order receipt time, etc. When the first stream data is a user login record, the fields of the first stream data include one or more of user name, password, name, role / authority, email address, mobile phone number, department / organization, login time, login Internet Protocol address (IP address for short), etc.
[0096] The first streaming data includes multiple data records. For example, the first streaming data is meteorological observation data, and the multiple data records are one or more of atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, rainfall, etc. collected by the first client at different times.
[0097] S402: The first client sends first streaming data to the primary key generation system. Correspondingly, the primary key generation system receives the first streaming data sent by the first client.
[0098] S403: The second client obtains the second streaming data.
[0099] The second streaming data is a data stream acquired by the second client in a time period. The second streaming data may be meteorological observation data, employee punch-in records, user login records, commodity transaction records, etc. This application does not limit the second streaming data. The second stream data includes multiple fields. When the second stream data is meteorological observation data, the fields of the second stream data include one or more of atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, and rainfall, etc. When the second stream data is an employee punch-in record, the fields of the second stream data include one or more of employee work number, employee name, punch-in time, punch-in location, punch-in method, punch-in type, overtime information, abnormal situation, etc.; when the second stream data is an employee punch-in record, the fields of the second stream data include one or more of purchase account name, purchase product type, purchase product quantity, paid amount, order number, order creation time, order payment time, order shipment time, and order receipt time, etc.; when the second stream data is a user login record, the fields of the second stream data include one or more of user name, password, name, role / authority, email address, mobile phone number, department / organization, login time, login Internet Protocol address (IP address for short), etc.
[0100] The second streaming data includes multiple data records. For example, the second streaming data is meteorological observation data, and the multiple data records are one or more of atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, rainfall, etc. collected by the second client at different times.
[0101] The fields of the second stream data and the first stream data are the same, that is, the first stream data and the second stream data are data of the same business type. For example, assuming that the first stream data and the second stream data both include the following fields: purchase account name, purchase product type, purchase product quantity, paid amount, order number, order creation time, order payment time, order shipment time and order receipt time, then the first stream data and the second stream data are both commodity transaction records. For another example, the first stream data and the second stream data both include the following fields: atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction, rainfall and ultraviolet index, then the first stream data and the second stream data are both meteorological observation information. However, the order of the fields can be the same or different. For example, both the first streaming data and the second streaming data include fields such as atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction and rainfall. However, the order of the fields in the first streaming data is atmospheric temperature, atmospheric humidity, air pressure, wind speed, wind direction and rainfall, while the order of multiple fields in the second streaming data is atmospheric humidity, atmospheric temperature, wind speed, air pressure, wind direction and rainfall.
[0102] S404: The second client sends second streaming data to the primary key generation system. Correspondingly, the primary key generation system receives the first streaming data sent by the second client.
[0103] S405: The primary key generation system generates a unified primary key for the multiple data records in the first streaming data and the multiple data records in the second streaming data.
[0104] The unified primary key includes multiple values, each value is a unique identifier of multiple data records in the first streaming data and multiple data records in the second streaming data. Each value of the unified primary key is generated based on the value of the corresponding field in the same unique field or combination of fields in the first streaming data and the second streaming data. Wherein, each value of the unified primary key is generated based on the value of the corresponding field in the same field or combination of fields in the first streaming data and the second streaming data in the following manner: the value in the unified primary key can be the value of the field or combination of fields with unique values in the first streaming data and the second streaming data, or, it is a hash value obtained by hashing the value of the field or combination of fields with unique values in the first streaming data and the second streaming data. The meaning of a unique field or combination of fields is: taking field A in the first streaming data and the second streaming data as an example, there is no duplicate value in the value of field A in the first streaming data and the second streaming data.
[0105] For the specific implementation process of S405, please refer to the relevant descriptions of Figures 5 and 7 below.
[0106] S406: The primary key generation system adds the unified primary key generated for each data record to each data record.
[0107] Adding a unified primary key to each data record means adding a unified primary key field carrying the corresponding key value to each data record.
[0108] Furthermore, the unified primary key can be added to the front of the first field of multiple fields in each data record, or to the back of the last field of multiple fields, or between any two fields in multiple fields. This application does not make specific restrictions on this.
[0109] S407: The primary key generation system sends each data record carrying the unified primary key to the data integration system. Correspondingly, the data integration system receives each data record carrying the unified primary key sent by the primary key generation system.
[0110] S408: The data integration system integrates each data record based on the unified primary key carried by each data record.
[0111] Taking the unified primary key field carried by each data record as A, when the data integration system receives multiple data records in the first streaming data and the first data record in the multiple data records in the second streaming data, it can establish a data table (or partition) corresponding to the primary key field A based on the primary key field A, and store the first data record in the data table (or partition) corresponding to the primary key field A. When the data integration system subsequently receives other data records carrying the primary key field A, it can find the corresponding data table (or partition) based on the primary key field A carried by other data records, and then store the other data records in the data table (or partition) corresponding to the primary key field A.
[0112] In a specific embodiment of the present application, when the data integration system subsequently receives other data records, it can also compare the primary key field and primary key value carried by the latest received data record with the primary key field A and primary key value carried by the already received data record to see whether they are the same. If it is determined that the primary key field and primary key value are the same, the latest received data record is determined to be a duplicate data record and can be discarded or otherwise processed. Otherwise, it is determined that the latest received data record is not a duplicate data record, and the data record is integrated into the primary key field A to find the corresponding data table (or partition).
[0113] In some possible embodiments, the data integration system may also use the value of the unified primary key of each data record as the index of the data record when integrating each data record, so that the data record can be subsequently located and searched based on the index.
[0114] Next, with reference to FIG5 , a primary key generation system is described to implement a process of generating a unified primary key for a data record set (referring to multiple data records in the first stream data and multiple data records in the second stream data). As shown in FIG5 , the process includes the following steps:
[0115] S501: The primary key generation system samples a data record set, and determines fields and / or field combinations with duplicate values in the sampled data records as first non-candidate keys.
[0116] The first non-candidate key is a field or combination of fields that cannot be used to uniquely identify the sampled data record due to duplicate values in the values corresponding to the first non-candidate key in the sampled data record. It is understood that the first non-candidate key cannot be used to uniquely identify the sampled data record because the sampled data record is derived from a set of data records. Therefore, the first non-candidate key cannot be used to uniquely identify the set of data records.
[0117] It is understood that when the sampled data records are different, the first non-candidate key determined based on the sampled data records is generally also different. The following is an explanation with reference to the meteorological observation information table shown in Table 1. It should be noted that the data in the meteorological observation information table shown in Figure 1 is only for example and is not to be considered as a specific limitation.
[0118] Table 1 Meteorological observation information
[0119] Take the data record set including the five rows of data records shown in Table 1 as an example:
[0120] Example 1. Assume that the data records sampled by the primary key generation system are the first row of data, the second row of data, and the fourth row of data in Table 1. The primary key generation system can determine that there is duplicate data "15.0" in the atmospheric temperature field, duplicate data "25.6" in the atmospheric humidity field, and duplicate data "15.0+25.6" in the "atmospheric temperature + atmospheric humidity" field combination by comparing the first row of data, the second row of data, and the fourth row of data. Therefore, the primary key generation system can determine that the fields with the same data are the atmospheric temperature field, the atmospheric humidity field, and the "atmospheric temperature + atmospheric humidity" field combination. The primary key generation system then determines the atmospheric temperature field, the atmospheric humidity field, and the "atmospheric temperature + atmospheric humidity" field combination as the three first non-candidate keys.
[0121] Example 2. Assume that the data records sampled by the primary key generation system are the first row of data, the second row of data, the fourth row of data, and the fifth row of data in Table 1. The primary key generation system can determine that there is duplicate data "14.5" in the atmospheric temperature field, duplicate data "25.6" in the atmospheric humidity field, duplicate data "0.11" in the wind speed field, duplicate data "15.0+25.6" in the "atmospheric temperature + atmospheric humidity" field combination, and duplicate data "14.5+0.11" in the "atmospheric temperature + wind speed" field combination by comparing the first row of data, the second row of data, the fourth row of data, and the fifth row of data. Therefore, the primary key generation system can determine that the fields with the same data are the atmospheric temperature field, the atmospheric humidity field, the wind speed field, the "atmospheric temperature + atmospheric humidity" field combination, and the "atmospheric temperature + wind speed" field combination. The primary key generation system then determines the atmospheric temperature field, the atmospheric humidity field, the wind speed field, the "atmospheric temperature + atmospheric humidity" field combination, and the "atmospheric temperature + wind speed" field combination as the five first non-candidate keys.
[0122] Based on the above examples, it can be seen that when the sampled data records are different, the first non-candidate keys determined based on the sampled data records are also different.
[0123] In this application, since the data volume of a data record set is usually relatively large, such as 5,000 or 10,000 data records, the primary key generation system samples the data record set and determines the first non-candidate key based on the sampled data records, rather than determining the first non-candidate key based on all data records. This can improve the efficiency of determining the first non-candidate key, thereby improving the efficiency of generating a unified primary key.
[0124] S502: The primary key generation system determines the remaining elements in the full set except the first non-candidate key as multiple first possible candidate keys, where the full set includes multiple fields and a combination of at least two of the multiple fields.
[0125] Continuing with the example of the four fields of observation time, atmospheric temperature, atmospheric humidity, and wind speed shown in Table 1, the full set includes: observation time field, atmospheric temperature field, atmospheric humidity field, wind speed field, "observation time + atmospheric temperature" field combination, "observation time + atmospheric humidity" field combination, "observation time + wind speed" field combination, "atmospheric temperature + atmospheric humidity" field combination, "atmospheric temperature + wind speed" field combination, "atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity" field combination, "observation time + atmospheric temperature + wind speed" field combination, "observation time + atmospheric humidity + wind speed" field combination, "atmospheric temperature + atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination, "atmospheric temperature + atmospheric humidity + wind speed" field combination, and "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination.
[0126] As can be seen from the description in S501, when the sampled data records are different, the first non-candidate keys determined based on the sampled data records are also different. It can be understood that when the first non-candidate keys are different, the multiple first possible candidate keys obtained by filtering the entire set based on the first non-candidate keys are usually also different. This is explained below with reference to Examples 1 and 2 in S501.
[0127] Example (1): First, take the atmospheric temperature field, atmospheric humidity field and the "atmospheric temperature + atmospheric humidity" field combination described in Example 1 in S501 as an example. The multiple first possible candidate keys determined by the primary key generation system include: observation time field, wind speed field, "observation time + atmospheric temperature" field combination, "observation time + atmospheric humidity" field combination, "observation time + wind speed" field combination, "atmospheric temperature + wind speed" field combination, "atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity" field combination, "observation time + atmospheric temperature + wind speed" field combination, "observation time + atmospheric temperature + wind speed" field combination, "observation time + atmospheric humidity + wind speed" field combination, "atmospheric temperature + atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination.
[0128] Example (2): Taking the atmospheric temperature field, atmospheric humidity field, wind speed field, “atmospheric temperature + atmospheric humidity” field combination and “atmospheric temperature + wind speed” field combination described in Example 2 in S501 as an example, the multiple first possible candidate keys determined by the primary key generation system include: observation time field, “observation time + atmospheric temperature” field combination, “observation time + atmospheric humidity” field combination, “observation time + wind speed” field combination, “atmospheric humidity + wind speed” field combination, “observation time + atmospheric temperature + atmospheric humidity” field combination, “observation time + atmospheric temperature + wind speed” field combination, “observation time + atmospheric humidity + wind speed” field combination, “observation time + atmospheric humidity + wind speed” field combination, “atmospheric temperature + atmospheric humidity + wind speed” field combination, and “observation time + atmospheric temperature + atmospheric humidity + wind speed” field combination.
[0129] Based on the above examples, it can be seen that when the first non-candidate keys are different, the multiple first possible candidate keys obtained by filtering the entire set based on the first non-candidate keys are also different.
[0130] S503: The primary key generation system verifies whether multiple first possible candidate keys can all be used as candidate keys based on the data record set. If it is determined that multiple first possible candidate keys can all be used as candidate keys, execute S504. If it is determined that multiple first possible candidate keys cannot all be used as candidate keys, execute S505-S507.
[0131] A candidate key is a field or combination of fields that can be used to uniquely identify a set of data records. That is, there are no duplicate values in the values corresponding to the candidate key in the data record set. Therefore, each value corresponding to the candidate key in the data record set can uniquely identify the data record to which each value belongs.
[0132] Continuing with Table 1 as an example, it can be seen from Table 1 that there are no duplicate values in the values corresponding to the observation time field, the "observation time + atmospheric temperature" field combination, the "observation time + atmospheric humidity" field combination, and the "observation time + atmospheric temperature + atmospheric humidity" field combination. Therefore, these fields and field combinations can be used to uniquely identify the data records in Table 1 and are all candidate keys.
[0133] As can be seen from the relevant description in S502, when the first non-candidate keys are different, the full set is filtered based on the first non-candidate keys, and the multiple first possible candidate keys obtained are also different. The multiple first possible candidate keys may all be candidate keys, or may not all be candidate keys. Therefore, it is necessary to verify whether the multiple first possible candidate keys can all be candidate keys. When it is determined that the multiple first possible candidate keys can all be used as candidate keys, S504 is executed, so that the first target candidate key selected from the multiple first possible candidate keys can be used to uniquely identify the data record set, so that the value of the unified primary key generated for each data record based on the value of the corresponding first target candidate key in each data record in the data record set can be used to uniquely identify each data record, thereby ensuring the accuracy of the unified primary key. When it is determined that there are non-candidate keys among the multiple first possible candidate keys, S505 and S506 are executed to sample more data records to filter out more or even all non-candidate keys from the full set to ensure the accuracy of the subsequently generated unified primary key.
[0134] In one possible embodiment, the primary key generation system may verify whether all of the multiple first possible candidate keys can be used as candidate keys if the number of the multiple first possible candidate keys reaches a certain threshold. If the number of the multiple first possible candidate keys does not reach the certain threshold, the primary key generation system may not verify whether all of the multiple first possible candidate keys can be used as candidate keys. The threshold value can be customized based on the actual scenario, such as being set to 3 or 5, and this application does not impose any specific restrictions on this.
[0135] When determining and verifying whether multiple first possible candidate keys can all be used as candidate keys, the i-th first possible candidate key among the multiple first possible candidate keys can be used as an example to illustrate how to perform the verification: the primary key generation system can determine whether there are duplicate values in the values corresponding to the i-th first possible candidate key in the data record set. If it is determined that there are duplicate values, it is determined that the i-th first possible candidate key cannot be used as a candidate key. If it is determined that there are no duplicate values, it is determined that the i-th first possible candidate key can be used as a candidate key.
[0136] The following describes the process of verifying the first possible candidate key in combination with example (1) and example (2) in S502.
[0137] First, take the multiple first possible candidate keys described in example (1) in S502 as an example, including: observation time field, wind speed field, "observation time + atmospheric temperature" field combination, "observation time + atmospheric humidity" field combination, "observation time + wind speed" field combination, "atmospheric temperature + wind speed" field combination, "atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity" field combination, "observation time + atmospheric temperature + wind speed" field combination, "observation time + atmospheric temperature + wind speed" field combination, "atmospheric temperature + atmospheric humidity + wind speed" field combination, and "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination as an example. Since the wind speed field has a repeated value "0.11" and the "atmospheric temperature + wind speed" field combination has a repeated value "14.5 + 0.11", while the remaining fields and field combinations do not have repeated values, the primary key generation system can determine that the multiple first possible candidate keys cannot all be candidate keys.
[0138] Taking the multiple first possible candidate keys described in example (2) in S502 as an example, including: observation time field, "observation time + atmospheric temperature" field combination, "observation time + atmospheric humidity" field combination, "observation time + wind speed" field combination, "atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity" field combination, "observation time + atmospheric temperature + wind speed" field combination, "observation time + atmospheric humidity + wind speed" field combination, "atmospheric temperature + atmospheric humidity + wind speed" field combination, "observation time + atmospheric temperature + atmospheric humidity + wind speed" field combination, as an example, since there are no duplicate values in these fields and field combinations, the primary key generation system can determine that multiple first possible candidate keys can all be used as candidate keys.
[0139] S504: The primary key generation system selects a first target candidate key from multiple first possible candidate keys, and generates a unified primary key value for each data record based on the value corresponding to the first target candidate key in each data record.
[0140] The primary key generation system selects the first target candidate key from multiple first possible candidate keys in at least the following ways:
[0141] Possible selection method (1): Select any one of multiple first possible candidate keys as the first target candidate key.
[0142] Possible selection method (2): The first possible candidate key with the most uniform data distribution among multiple first possible candidate keys is selected as the first target candidate key, wherein the data distribution of the first possible candidate key is all the data corresponding to the first possible candidate key in the data record set.
[0143] Possible selection method (3): select the first possible candidate key with uniform data distribution and the least number of fields among multiple first possible candidate keys as the first target candidate key.
[0144] In a specific implementation, the first possible candidate key with uniform (or most uniform) data distribution can be selected in the following manner: by drawing a frequency distribution histogram of the data of multiple first possible candidate keys, observing the distribution of the data in different intervals, and thus finding one from the multiple first possible candidate keys as the first target candidate key. Optionally, statistical indicators can also be used to judge the uniformity of the data, such as the mean, median and mode. The mean is used to understand the overall mean of the data. If the mean of the data is not much different from the median and mode, it means that the data is relatively uniform. Another example is the variance and standard deviation. The variance and standard deviation reflect the degree of dispersion of the data. If the variance or standard deviation is small, the data is relatively uniform. If the variance or standard deviation is large, the data is relatively non-uniform. This application does not limit the method of judging the uniform distribution of multiple first possible candidate keys.
[0145] Taking the use of frequency distribution histograms to determine the uniformity of data distribution for multiple first possible candidate keys as an example, assuming that the two first possible candidate keys correspond to the atmospheric temperature field and the atmospheric humidity field, each corresponding to 6000 values, the histogram of the atmospheric temperature field is shown in FIG6 (a), and the histogram of the atmospheric humidity field is shown in FIG6 (b). Then, the primary key generation system can determine the data distribution of the atmospheric temperature field through these two histograms as follows: there are 1000 values less than 29°C, 1500 values between 29-30°C, 2800 values between 30-31°C, and 700 values greater than 31°C. It can also determine the data distribution of the atmospheric humidity field as follows: there are 1500 values less than 25%, 1600 values between 25%-26%, 1800 values between 26%-27%, and 1100 values greater than 27%. Since the data distribution of the atmospheric humidity field is more uniform than that of the atmospheric temperature field, the primary key generation system prefers the atmospheric humidity field as the first target candidate key.
[0146] Through the above possible selection methods (2) / (3), since the data distribution of the first target candidate key is uniform, it can be understood that this can make the data distribution of the unified primary key generated for the data record set based on the value of the first target candidate key also relatively uniform. Since in the database, the primary key is usually used for indexing to speed up data retrieval, if the data distribution of the primary key is uniform, then the index effect will be better and the query speed will be faster.
[0147] It should be understood that the above-mentioned possible selection methods (1) to (3) are merely examples of selection methods for the primary key generation system to select the first target candidate key from multiple first possible candidate keys. In specific implementations, other methods can also be referred to for selection. For example, the primary key generation system provides multiple first possible candidate keys to the user, and then obtains the first possible candidate key selected by the user from the multiple first possible candidate keys as the first target candidate key. This application does not specifically limit the selection method.
[0148] In a possible embodiment, the primary key generation system may hash the value corresponding to the first target candidate key in each data record, and use the obtained hash value as the value of the unified primary key of each data record.
[0149] In another possible embodiment, the primary key generation system may use the value corresponding to the first target candidate key in each data record as the value of the unified primary key of each data record.
[0150] It can be understood that the primary key generation system hashes the value corresponding to the first target candidate key in each data record and uses the obtained hash value as the value of the unified primary key of each data record. Compared with using the value corresponding to the first target candidate key in each data record as the value of the unified primary key of each data record, since the hash function is a one-way function, that is, the original value cannot be deduced from the hash value, it can avoid direct exposure of the value corresponding to the first target candidate key in each data record, thereby improving data security.
[0151] Furthermore, because hash functions typically map input data uniformly to the output space, this ensures that the generated hash values are relatively evenly distributed within the unified primary key. Uniformly distributed primary keys help improve data query performance and indexing efficiency. Furthermore, because hash values generated by hash functions typically have a fixed length and are unaffected by the length of the input data, this ensures that the unified primary key value occupies a fixed amount of space during storage and indexing, improving database storage efficiency and query performance. Furthermore, because the hash value is calculated based on the value of the first target candidate key in each data record and is independent of the value of the first target candidate key in each data record, the generated hash value remains unchanged even if the value of the first target candidate key in each subsequent data record changes. Therefore, the value of the primary key field remains unchanged. This is very useful for data updates and maintenance, as it avoids primary key conflicts and data consistency issues caused by changes to the original data. In short, using hash values as unified primary key values offers advantages over directly using the original data as unified primary key values, including data protection, uniqueness assurance, uniform distribution, fixed length, and irrelevance. These advantages can improve database security, performance, and maintainability.
[0152] S505: The primary key generation system continues to sample the data record set, and determines the fields and / or field combinations with duplicate values in the continuously sampled data records as the second non-candidate key.
[0153] The second non-candidate key is a field or combination of fields that cannot be used to uniquely identify the data records being sampled further due to duplicate values in the values corresponding to the second non-candidate key in the data records being sampled further. It is understood that the second non-candidate key cannot be used to uniquely identify the data records being sampled further because the data records being sampled further are derived from the data record set. Therefore, the second non-candidate key cannot be used to uniquely identify the data record set.
[0154] S506: The primary key generation system determines the remaining elements in the entire set except the first non-candidate key and the second non-candidate key as multiple second possible candidate keys.
[0155] For an introduction to the entire set, please refer to the introduction in S502.
[0156] S507: The primary key generation system verifies whether multiple second possible candidate keys can all be used as candidate keys based on the data record set. If it is determined that multiple second possible candidate keys can all be used as candidate keys, execute S508. If it is determined that multiple second possible candidate keys cannot all be used as candidate keys, execute S509.
[0157] S508: The primary key generation system selects a second target candidate key from the plurality of second possible candidate keys, and generates a unified primary key value for each data record based on the value corresponding to the second target candidate key in each data record.
[0158] In a possible embodiment, the primary key generation system may hash the value corresponding to the second target candidate key in each data record, and use the obtained hash value as the value of the unified primary key of each data record.
[0159] In another possible embodiment, the primary key generation system may use the value corresponding to the second target candidate key in each data record as the value of the unified primary key of each data record.
[0160] S509: The primary key generation system continues to sample the data record set, and determines the fields and / or field combinations with duplicate values in the continuously sampled data records as the third non-candidate key.
[0161] The specific implementation process of S505-S509 is similar to the specific implementation process of S501-S505. Please refer to the relevant description of S501-S505. For the sake of brevity of the description, it will not be elaborated again.
[0162] It can be seen that based on the embodiment of Figure 5, by analogy, the primary key generation system can filter out all non-candidate keys from the entire set, leaving all candidate keys, and select suitable candidate keys from the remaining candidate keys, and generate a unified primary key value for each data record based on the suitable candidate keys.
[0163] In a possible embodiment, when the primary key generation system selects the first target candidate key / second target candidate key, when obtaining a new data record including the same multiple fields, it can hash the value corresponding to the first target candidate key / second target candidate key in the new data record, and use the obtained hash value as the value of the unified primary key of the new data record and add it to the new data record, or use the value corresponding to the first target candidate key in the new data record as the value of the unified primary key of the new data record and add it to the new data record, and then send the new data record carrying the value of the corresponding unified primary key to the data integration system for integration by the data integration system.
[0164] In a possible embodiment, after selecting the first target candidate key / second target candidate key, the primary key generation system can also send a notification carrying the first target candidate key / second target candidate key to the client (including the first client and the second client), instructing the client to hash the value corresponding to the first target candidate key / second target candidate key in the new data record when obtaining a new data record including the same multiple fields, and use the obtained hash value as the value of the unified primary key of the new data record and add it to the new data record, or use the value corresponding to the first target candidate key / second target candidate key in the new data record as the value of the unified primary key of the new data record and add it to the new data record, and then send the new data record carrying the value of the corresponding unified primary key to the data integration system for integration by the data integration system.
[0165] Next, with reference to FIG7 , another implementation process of the primary key generation system generating a unified primary key for a data record set (referring to multiple data records in the first stream data and multiple data records in the second stream data) is described. As shown in FIG7 , the process includes the following steps:
[0166] S701: The primary key generation system samples a data record set, and determines fields and / or field combinations with duplicate values in the sampled data records as first non-candidate keys.
[0167] S701 is the same as S501. Please refer to the description of S501.
[0168] S702: The primary key generation system determines the remaining elements in the full set except the first non-candidate key as multiple first possible candidate keys, where the full set includes multiple fields and a combination of at least two of the multiple fields.
[0169] S702 is the same as S502. Please refer to the description of S502.
[0170] S703: The primary key generation system verifies whether there is at least one first possible candidate key that can be used as a candidate key among multiple first possible candidate keys based on the data record set. If it is determined that there is at least one first possible candidate key that can be used as a candidate key, S704 is executed. If it is determined that there is not at least one first possible candidate key that can be used as a candidate key, S705 is executed.
[0171] For an introduction to candidate keys, please refer to the introduction to candidate keys in S503.
[0172] The specific process of the primary key generation system verifying whether there is at least one first possible candidate key that can be used as a candidate key among multiple first possible candidate keys based on a data record set is similar to the specific process of the primary key generation system verifying whether multiple first possible candidate keys can all be used as candidate keys based on a data record set described in S503. Please refer to the relevant description in S503.
[0173] S704: The primary key generation system selects a first target candidate key from at least one first possible candidate key, and generates a unified primary key value for each data record based on the value corresponding to the first target candidate key in each data record.
[0174] S704 is similar to S504. Please refer to the description of S504.
[0175] S705: The primary key generation system sends a notification to the client, notifying the user of the client that a unified primary key cannot be generated.
[0176] The following describes the entire process of streaming data integration in the system shown in Figure 2 with a more detailed example. Referring to Figure 8 , Figure 8 is a schematic diagram of another streaming data integration process provided by an embodiment of the present application. The process from left to right in Figure 8 is as follows:
[0177] First, the clients 100 on multiple source ends transmit the data records generated by the source ends to the primary key generation system 200 in real time.
[0178] S1, the primary key generation system 200 determines whether each received data record carries a valid primary key. If it is determined that the data record carries a valid primary key, the data record carrying the valid primary key is sent to the data integration system 300 for integration. If it is determined that the data record does not carry a valid primary key, S2 is executed.
[0179] Specifically, the primary key generation system 200 can pre-store multiple primary key fields of the data integrated in the data integration system 300. When receiving each data record, it determines whether the data record carries a primary key field. If it is determined that the data record does not carry a primary key field, it is determined that the data record does not carry a valid primary key. If it is determined that the data record carries a primary key field, the primary key field carried by the data record is matched with multiple primary key fields pre-stored locally. If it is determined that there is a primary key field in the multiple primary key fields that is the same as the primary key field carried by the data record, it is determined that the data record carries a valid primary key. Otherwise, it is determined that the data record does not carry a valid primary key. Here, the multiple primary key fields pre-stored locally by the primary key generation system 200 can be the primary key fields of the corresponding data tables (or partitions) that have been established in the data integration system 300.
[0180] S2, the primary key generation system 200 determines whether a unified primary key can be generated for the data records that do not carry a valid primary key based on the initial broadcast variables. If it is determined that a unified primary key cannot be generated for the data records that do not carry a valid primary key, the data records that do not carry a valid primary key are stored in the streaming database and S3-S6 are executed. If it is determined that a unified primary key can be generated for the data records that do not carry a valid primary key, S7-S9 are executed.
[0181] When the primary key generation system 200 is initialized, it generates an initial broadcast variable, which indicates that a high-confidence candidate key has not yet been mined (the high-confidence candidate key is the first target candidate key described in the embodiment of FIG. 4 ).
[0182] Specifically, the primary key generation system 200 can determine whether the initial broadcast variable tells a high-confidence candidate key. If it is determined that the initial broadcast variable does not tell a high-confidence candidate key, it is determined that a unified primary key cannot be generated for data records that do not carry a valid primary key. Otherwise, it is determined that a unified primary key can be generated for data records that do not carry a valid primary key.
[0183] S3, the primary key generation system 200 samples data records in the streaming database.
[0184] S4, the primary key generation system 200 performs high-confidence candidate key mining based on the sampled data records. If a high-confidence candidate key is mined, S5 is executed. If no high-confidence candidate key is mined, S3 is executed again, more data records in the sampling streaming database are sampled, and S4 is executed based on the continued sampled data records, and so on.
[0185] S5: The primary key generation system 200 writes the high-confidence candidate key into the first broadcast variable.
[0186] S6: The primary key generation system 200 updates the broadcast variable, that is, uses the first broadcast variable to update the initial broadcast variable.
[0187] It can be understood that after the primary key generation system 200 updates the broadcast variable, in S2, it can determine whether a unified primary key can be generated for the newly received data record that does not carry a valid primary key based on the first broadcast variable.
[0188] S7, the primary key generation system 200 extracts the value corresponding to the high-confidence candidate key in the data record, and generates a unified primary key value for the data record based on the extracted value.
[0189] S8, the primary key generation system 200 adds the unified primary key (including the value of the unified primary key) to the data record and sends it to the data integration system 300 for integration.
[0190] In this example, when the primary key generation system 200 mines a high-confidence candidate key, as shown by the dotted arrow in Figure 8, the primary key generation system 200 can also execute S7-S8 on each data record in the streaming database, generate a unified primary key value for each data record, and then add the unified primary key (including the unified primary key value) to the data record and send it to the data integration system 300 for integration.
[0191] The process of the above-mentioned primary key generation system 200 performing high-confidence candidate key mining based on the sampled data records can be referred to S501-S509 in the embodiment of Figure 5, in which the primary key generation system determines the target candidate key (first target candidate key, second target candidate key...) based on the sampled data records. The process of the primary key generation system 200 generating a unified primary key value for each data record can be referred to the relevant description in S504. For the sake of brevity of the specification, it will not be elaborated here.
[0192] The process of the data integration system 300 integrating multiple data records based on the unified primary key can be referred to the relevant description of S406 above.
[0193] It can be seen that through the above embodiments, the primary key generation system 200 can mine high-confidence candidate keys based on multiple data records that do not carry valid primary keys, and then generate a unified primary key for multiple data records based on the mined high-confidence candidate keys, thereby facilitating the data integration system 300 to integrate multiple data records based on the unified primary key.
[0194] In order to facilitate a clearer understanding of the specific process of generating a unified primary key for streaming data and integrating streaming data based on the unified primary key in the architecture shown in Figure 3, a more detailed introduction is given below in conjunction with a schematic diagram of a streaming data integration process provided in an embodiment of the present application shown in Figure 9.
[0195] It should be noted that in FIG9 , FIG3 is taken as an example to illustrate the method including two source terminals. As shown in FIG9 , the method includes the following steps:
[0196] S901: The client receives first streaming data sent by a first source.
[0197] S902: The client receives second streaming data sent by the second source.
[0198] S903: The client sends a primary key generation request to the primary key generation system. Correspondingly, the primary key generation system receives the primary key generation request sent by the client.
[0199] The primary key generation request carries the first streaming data and the second streaming data.
[0200] S904: The primary key generation system generates a unified primary key for the data record set.
[0201] S905: The primary key generation system adds the unified primary key generated for each data record to each data record.
[0202] S906: The primary key generation system sends each data record carrying the unified primary key to the data integration system. Correspondingly, the data integration system receives each data record carrying the unified primary key sent by the primary key generation system.
[0203] S907: The data integration system integrates each data record based on the unified primary key carried by each data record.
[0204] For the concepts of the first streaming data, the second streaming data, the data record set, the unified primary key, etc. in FIG9 , please refer to the description of the related concepts in the embodiment of FIG4 , and for the sake of brevity of the specification, they will not be elaborated again.
[0205] It can be seen from the steps shown in FIG4 that the steps shown in FIG9 are the same as or similar to the steps shown in FIG4 . Please refer to the specific description of the steps shown in FIG4 . For the sake of brevity of the specification, they will not be described in detail.
[0206] It should be noted that the embodiments of Figures 4 and 9 are described by taking the example of generating a unified primary key for two pieces of streaming data. The inventive concept provided in this application can also be applied to generating a unified primary key for one or three or more pieces of streaming data. The specific implementation process is similar to that of the embodiments of Figures 4 and 9. For the sake of brevity of the specification, it will not be elaborated again.
[0207] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0208] The above describes in detail the primary key generation system and primary key determination method provided in the embodiments of the present application. Based on the same inventive concept, the following further introduces the primary key determination device and computing device cluster provided in the embodiments of the present application.
[0209] The primary key determination device can be applied to the aforementioned primary key generation system to implement the various functions of the primary key generation system. The primary key determination device can be divided into one or more unit modules. For example, see Figure 10 , which is a schematic diagram of the structure of a primary key determination device provided in an embodiment of the present application. As shown in Figure 10 , the primary key determination device includes an acquisition module 1010, a processing module 1030, and a sampling module 1020. The functions of each module of the primary key determination device 1000 are exemplarily described below.
[0210] The acquisition module 1010 is configured to acquire a plurality of data records, where the plurality of data records include the same plurality of fields.
[0211] The sampling module 1020 is configured to sample multiple data records.
[0212] The processing module 1030 is configured to determine fields and / or field combinations with duplicate values in the sampled data records as first non-candidate keys.
[0213] The processing module 1030 is configured to determine the remaining elements in the full set except the first non-candidate key as multiple first possible candidate keys, where the full set includes multiple fields and a combination of at least two of the multiple fields.
[0214] The processing module 1030 is also used to select a first target candidate key from multiple first possible candidate keys when it is determined based on multiple data records that multiple first possible candidate keys can all be used as candidate keys, or to select a first target candidate key from at least one first possible candidate key when it is determined based on multiple data records that at least one first possible candidate key among multiple first possible candidate keys can be used as a candidate key.
[0215] The processing module 1030 is further configured to generate a unified primary key value for each data record based on the value corresponding to the first target candidate key in each data record.
[0216] In some possible embodiments, none of the multiple data records carry a primary key, or the multiple data records carry different primary keys, or some of the multiple data records carry a primary key and the rest do not.
[0217] In some possible embodiments, the sampling module 1020 is further configured to, when it is determined based on the multiple data records that the multiple first possible candidate keys cannot all be candidate keys, continue sampling the multiple data records, and the processing module 1030 is further configured to perform the following operations:
[0218] The fields and / or field combinations with repeated values in the data records that will be continuously sampled are determined as second non-candidate keys; the remaining elements in the full set except the first non-candidate key and the second non-candidate key are determined as multiple second possible candidate keys; when it is determined based on multiple data records that multiple second possible candidate keys can all be used as candidate keys, a second target candidate key is selected from the multiple second possible candidate keys; based on the value corresponding to the second target candidate key in each data record, a unified primary key value for each data record is generated.
[0219] In some possible embodiments, the processing module 1030 is configured to select a first possible candidate key with the most uniform data distribution among the multiple first possible candidate keys as the first target candidate key.
[0220] In some possible embodiments, the processing module 1030 is configured to hash the value corresponding to the first target candidate key in each data record, and use the obtained hash value as the value of the unified primary key of each data record.
[0221] In some possible embodiments, the processing module 1030 is configured to use the value corresponding to the first target candidate key in each data record as the value of the unified primary key of each data record.
[0222] In some possible embodiments, the acquisition module 1010 is also used to acquire a new data record, wherein the new data record also includes multiple fields; the processing module 1030 is also used to generate a unified primary key value of the new data record based on the value corresponding to the first target candidate key in the new data record.
[0223] In some possible embodiments, as shown in Figure 10, the device 1000 also includes a sending module 1040; the above-mentioned processing module 1030 is also used to add the value of the unified primary key of each data record to each data record; the above-mentioned sending module 1040 is used to send each data record to the data integration system for integration.
[0224] In some possible embodiments, the multiple data records belong to the same data stream, or belong to different data streams.
[0225] In some possible embodiments, when multiple data records belong to different data streams, the different data streams originate from the same source end or different source ends.
[0226] In some possible embodiments, after multiple data records are obtained, the above-mentioned acquisition module 1010 can store the obtained multiple data records in the same data bucket (the first data bucket as shown in Figure 10), and the subsequent sampling module 1020 can sample multiple data records from the first data bucket, and the processing module 1030 determines multiple possible candidate keys (such as the above-mentioned multiple first possible candidate keys and multiple second possible candidate keys) based on the sampled data records, and stores the determined multiple possible candidate keys in another data bucket (the second data bucket as shown in Figure 10). In addition, the processing module 1030 can also store the data records with the value of the unified primary key added in another data bucket (the third data bucket as shown in Figure 10), and the subsequent sending module 1040 can obtain the data records with the value of the unified primary key added from the third data bucket and send them to the data integration system. Optionally, the aforementioned multiple data records, multiple possible candidate keys, and data records with added unified primary key values may all be stored in the same data bucket, or, multiple data records and multiple possible candidate keys may be stored in the same data bucket, and data records with added unified primary key values may be stored in another data bucket, or, multiple data records and data records with added unified primary key values may be stored in the same data bucket, and multiple possible candidate keys may be stored in another data bucket. This application does not make any specific restrictions on this.
[0227] In a specific implementation, the acquisition module 1010, sampling module 1020, processing module 1030, and sending module 1040 can all be implemented via software or hardware. For example, the implementation of processing module 1030 will be described below using processing module 1030 as an example. Similarly, the implementation of acquisition module 1010, sampling module 1020, and sending module 1040 can refer to the implementation of processing module 1030.
[0228] As an example of a software functional unit, the processing module 1030 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the processing module 1030 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0229] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0230] As an example of a hardware functional unit, processing module 1030 may include at least one computing device, such as a server. Alternatively, processing module 1030 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0231] When processing module 1030 includes multiple computing devices, the multiple computing devices can be distributed in the same region or in different regions. The multiple computing devices included in processing module 1030 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in processing module 1030 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0232] It should be noted that, in other embodiments, the processing module 1030 can be used to execute any step among the steps performed by the primary key generation system in Figures 4, 5, and 7-9, the acquisition module 1010 can be used to execute any step among the steps performed by the primary key generation system in Figures 4, 5, and 7-9, the sampling module 1020 can be used to execute any step among the steps performed by the primary key generation system in Figures 4, 5, and 7-9, and the sending module 1040 can be used to execute any step among the steps performed by the primary key generation system in Figures 4, 5, and 7-9. The steps that the acquisition module 1010, the processing module 1030, the sampling module 1020, and the sending module 1040 are responsible for implementing can be specified as needed. By using the acquisition module 1010, the processing module 1030, the sampling module 1020, and the sending module 1040, different steps among the steps performed by the primary key generation system in Figures 4, 5, and 7-9 are respectively implemented to realize all the functions of the primary key determination device 1000.
[0233] It should be understood that the functions of the modules described above are merely functions that the primary key determination device 1000 may have in some embodiments of the present application, and the present application does not limit the functions of the modules.
[0234] It should also be understood that Figure 10 is an exemplary division method, and the primary key determination device 1000 can also include more or fewer modules. The division method of the modules in the primary key determination device 1000 can be flexibly adjusted based on the actual business scenario, and this application does not make specific limitations.
[0235] An embodiment of the present application also provides a computing device 1100, which can deploy the primary key determination device 1000 shown in Figure 10, or it can be the aforementioned primary key generation system. The operations and / or functions of each module in the computing device 1100 are respectively to implement the corresponding steps performed by the primary key generation system in the methods shown in Figures 4, 5, and 7-9.
[0236] As shown in FIG. 11 , a computing device 1100 includes a processor 1110 , a memory 1120 , and a communication interface 1130 , wherein the processor 1110 , the memory 1120 , and the communication interface 1130 may be interconnected via a bus 1140 .
[0237] The processor 1110 can read the program code (including instructions) stored in the memory 1120 and execute the program code stored in the memory 1120, so that the computing device 1100 executes the steps performed by the primary key generation system in Figures 4, 5, 7-9, or causes the computing device 1100 to deploy the primary key determination device 1000.
[0238] Processor 1110 can be implemented in various forms, such as a central processing unit (CPU), or a combination of a CPU and a hardware chip. The hardware chip can be an ASIC, a programmable logic device (PLD), or a combination thereof. The PLD can be a CPLD, an FPGA, a GAL, or any combination thereof. Processor 1110 executes various types of digitally stored instructions, such as software or firmware programs stored in memory 1120, enabling computing device 1100 to provide a wide variety of services.
[0239] In a specific implementation, as an embodiment, the processor 1110 includes one or more CPUs.
[0240] In a specific implementation, as an embodiment, the computing device 1100 also includes multiple processors, each of which can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here refers to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0241] Memory 1120 is used to store program code, and is controlled by processor 1110 to execute the steps performed by the primary key generation system in Figures 4, 5, and 7-9. The program code may include one or more software modules, which may be the software modules provided in the embodiment of Figure 10, such as acquisition module 1010, processing module 1030, sampling module 1020, and sending module 1040.
[0242] The memory 1120 may include a volatile memory, such as a random access memory (RAM); the memory 1120 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); the memory 1120 may also include a combination of the above types.
[0243] The communication interface 1130 may be a wired interface (e.g., an Ethernet interface, a fiber optic interface, or other types of interfaces (e.g., an InfiniBand interface)) or a wireless interface (e.g., a cellular network interface or a wireless local area network interface) for communicating with other computing devices or apparatuses. The communication interface 1130 may utilize a protocol suite based on the Transmission Control Protocol / Internet Protocol (TCP / IP), such as the Remote Function Call (RFC) protocol, the Simple Object Access Protocol (SOAP) protocol, the Simple Network Management Protocol (SNMP) protocol, the Common Object Request Broker Architecture (CORBA) protocol, and distributed protocols.
[0244] The bus 1140 may be a Peripheral Component Interconnect Express (PCIe) bus, an Extended Industry Standard Architecture (EISA) bus, a unified bus (Ubus or UB), a Compute Express Link (CXL), a Cache Coherent Interconnect for Accelerators (CCIX), etc. The bus 1140 may be divided into an address bus, a data bus, a control bus, etc.
[0245] In addition to the data bus, bus 1140 may also include a power bus, a control bus, and a status signal bus. However, for clarity, various buses are labeled as bus 1140 in the figure. For ease of illustration, FIG11 uses only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0246] The computing device 1100 is used to execute the steps in FIG. 4 , FIG. 5 , and FIG. 7 - FIG. 9 performed by the primary key generation system. The specific implementation process is detailed in the method embodiment described above and will not be repeated here.
[0247] It should be understood that computing device 1100 is merely an example provided in the embodiments of the present application, and computing device 1100 may have more or fewer components than those shown in FIG11 , may combine two or more components, or may have different configurations of components. For matters not shown or described in the embodiments of the present application, please refer to the relevant descriptions in the embodiments of FIG1 to FIG10 , and no further description will be given here.
[0248] The present application also provides a computing device cluster 1200, which can deploy the primary key determination device 1000 shown in Figure 10, or the aforementioned primary key generation system 200. The operations and / or functions of each module in the computing device cluster 1200 are respectively to implement the corresponding steps performed by the primary key generation system in the methods shown in Figures 4, 5, and 7-9.
[0249] As shown in FIG12 , the computing device cluster 1200 includes at least one computing device 1100. The memory 1120 in one or more computing devices 1100 in the computing device cluster may store the same instructions for executing the steps performed by the primary key generation system in FIG4 , FIG5 , and FIG7 - FIG9 . The computing device 1100 may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device 1100 may also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0250] In some possible implementations, the memory 1120 of one or more computing devices 1100 in the computing device cluster 1200 may also store instructions for executing the steps performed by the primary key generation system in Figures 4, 5, and 7-9. In other words, the combination of one or more computing devices 1100 can collectively execute instructions for executing the steps performed by the primary key generation system in Figures 4, 5, and 7-9.
[0251] It should be noted that the memory 1120 in different computing devices 1100 in the computing device cluster 1200 may store different instructions, each for executing part of the functions of the primary key determination apparatus 1000. That is, the instructions stored in the memory 1120 in different computing devices 1100 may implement the functions of one or more of the acquisition module 1010, the processing module 1030, the sampling module 1020, and the sending module 1040.
[0252] In some possible implementations, one or more computing devices 1100 in the computing device cluster 1200 may be connected via a network. The network may be a wide area network or a local area network, etc. FIG13 shows a possible implementation. As shown in FIG13 , two computing devices 1100A and 1100B are connected via a network. Specifically, the connection to the network is made through a communication interface in each computing device. In this type of possible implementation, the memory 1120 in the computing device 1100A stores instructions for executing the functions of the acquisition module 1010 and the sending module 1040. At the same time, the memory 1120 in the computing device 1100B stores instructions for executing the functions of the processing module 1030 and the sampling module 1020.
[0253] The connection method between the computing device cluster 1200 shown in Figure 13 may be based on the consideration that the primary key determination method provided in the embodiment of the present application needs to sample a large number of data records, analyze the sampled data records, and generate a unified primary key for a large number of data records. Therefore, it is considered to hand over the functions implemented by the sampling module 1020 and the processing module 1030 to the computing device 1100B for execution.
[0254] It should be understood that the functionality of the computing device 1100A shown in FIG13 may also be implemented by multiple computing devices 1100. Similarly, the functionality of the computing device 1100B may also be implemented by multiple computing devices 1100.
[0255] The present application also provides a computer program product comprising instructions, which may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the computer program product causes the at least one computing device to execute the steps performed by the primary key generation system in Figures 4, 5, and 7-9.
[0256] The present application also provides a computer-readable storage medium, which can be any available medium that can be stored by a computing device or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a high-density digital video disc (DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the steps performed by the primary key generation system in Figures 4, 5, and 7-9.
[0257] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0258] In the above embodiments, it can be implemented in whole or in part by software, hardware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium, or a semiconductor medium.
[0259] The above is only a specific embodiment of the present application. Those skilled in the art may conceive of changes or substitutions based on the specific embodiment provided in this application, which should all be included in the scope of protection of this application.
Claims
1. A primary key determination method, characterized in that: The method comprises: The primary key generation system obtains a plurality of data records, wherein the plurality of data records include the same plurality of fields; Sampling the multiple data records, and determining fields and / or field combinations with duplicate values in the sampled data records as first non-candidate keys; Determine the remaining elements in the full set except the first non-candidate key as a plurality of first possible candidate keys, wherein the full set includes the plurality of fields and a combination of at least two of the plurality of fields; If it is determined based on the multiple data records that all of the multiple first possible candidate keys can be used as candidate keys, a first target candidate key is selected from the multiple first possible candidate keys; or, if it is determined based on the multiple data records that at least one of the multiple first possible candidate keys can be used as a candidate key, a first target candidate key is selected from the at least one first possible candidate key; Based on the value of the first target candidate key in each data record, a value of the unified primary key of each data record is generated.
2. The method according to claim 1, characterized in that None of the multiple data records carry a primary key, or the multiple data records carry different primary keys, or some of the multiple data records carry a primary key and the rest do not carry a primary key.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: If it is determined based on the multiple data records that the multiple first possible candidate keys cannot all be candidate keys, continue sampling the multiple data records, and determine fields and / or field combinations with repeated values in the data records to be sampled as second non-candidate keys; Determine the remaining elements in the full set except the first non-candidate key and the second non-candidate key as a plurality of second possible candidate keys; When it is determined based on the multiple data records that the multiple second possible candidate keys can all be used as the candidate key, selecting a second target candidate key from the multiple second possible candidate keys; Based on the value of the second target candidate key corresponding to each data record, the value of the second target candidate key of the unified primary key of each data record is generated.
4. The method according to any one of claims 1 to 3, characterized in that: The step of selecting a first target candidate key from the plurality of first possible candidate keys comprises: The first possible candidate key with the most uniform data distribution among the multiple first possible candidate keys is used as the first target candidate key.
5. The method according to any one of claims 1 to 4, characterized in that: The generating, based on the value of the first target candidate key in each data record, the value of the unified primary key of each data record includes: A hash is performed on the value corresponding to the first target candidate key in each data record, and the obtained hash value is used as the value of the unified primary key of each data record.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: When a new data record is received, a value of a unified primary key of the new data record is generated based on a value corresponding to the first target candidate key in the new data record, wherein the new data record also includes the multiple fields.
7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Adding the value of the unified primary key of each data record to each data record; Each of the data records is sent to a data integration system for integration.
8. The method according to any one of claims 1 to 7, characterized in that: The multiple data records belong to the same data stream, or belong to different data streams.
9. The method according to claim 8, characterized in that In the case that the multiple data records belong to different data streams, the different data streams originate from the same source end or different source ends.
10. A primary key determination device, characterized in that: The device comprises: An acquisition module, used for acquiring a plurality of data records, wherein the plurality of data records include a plurality of identical fields; A sampling module, configured to sample the plurality of data records; A processing module, configured to determine fields and / or field combinations having duplicate values in the sampled data records as first non-candidate keys; The processing module is used to determine the remaining elements in the full set except the first non-candidate key as a plurality of first possible candidate keys, wherein the full set includes the plurality of fields and a combination of at least two of the plurality of fields; The processing module is further configured to select a first target candidate key from the plurality of first possible candidate keys when it is determined based on the plurality of data records that all of the plurality of first possible candidate keys can be used as candidate keys, or select a first target candidate key from the at least one first possible candidate key when it is determined based on the plurality of data records that at least one of the plurality of first possible candidate keys can be used as a candidate key; The processing module is further configured to generate a value of a unified primary key for each data record based on a value corresponding to the first target candidate key in each data record.
11. The device according to claim 10, characterized in that None of the multiple data records carry a primary key, or the multiple data records carry different primary keys, or some of the multiple data records carry a primary key and the rest do not carry a primary key.
12. The device according to claim 10 or 11, characterized in that The sampling module is further configured to continue sampling the plurality of data records if it is determined based on the plurality of data records that the plurality of first possible candidate keys cannot all be candidate keys; The processing module is further used to determine the fields and / or field combinations with repeated values in the continuously sampled data records as the second non-candidate key; The processing module is further used to determine the remaining elements in the full set except the first non-candidate key and the second non-candidate key as a plurality of second possible candidate keys; The processing module is further configured to select a second target candidate key from the plurality of second possible candidate keys when it is determined based on the plurality of data records that the plurality of second possible candidate keys can all serve as the candidate key; The processing module is further configured to generate a value of a unified primary key for each data record based on a value corresponding to the second target candidate key in each data record.
13. The device according to any one of claims 10 to 12, characterized in that The processing module is used to use the first possible candidate key with the most uniform data distribution among the multiple first possible candidate keys as the first target candidate key.
14. The device according to any one of claims 10 to 13, characterized in that The processing module is used to hash the value corresponding to the first target candidate key in each data record, and use the obtained hash value as the value of the unified primary key of each data record.
15. The device according to any one of claims 10 to 14, characterized in that The acquisition module is further used to acquire a new data record, wherein the new data record also includes the multiple fields; The processing module is further configured to generate a value of a unified primary key of the new data record based on a value corresponding to the first target candidate key in the new data record.
16. The device according to any one of claims 10 to 15, characterized in that The device also includes: a sending module; The processing module is further used to add the value of the unified primary key of each data record to each data record; The sending module is used to send each data record to the data integration system for integration.
17. The device according to any one of claims 10 to 16, characterized in that The multiple data records belong to the same data stream, or belong to different data streams.
18. The device according to claim 17, characterized in that In the case that the multiple data records belong to different data streams, the different data streams originate from the same source end or different source ends.
19. A computing device cluster, characterized in that: It includes at least one computing device, each of which includes a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Sequence generation method and device, computer equipment and storage medium
CN110222048A
Primary key extraction method and device and storage medium
CN113761185A
Primary key generation method and electronic equipment
CN115794817A
Method for identifying primary key and foreign key in relational database table
CN117349346A