Data matching method, device, computer equipment and readable storage medium

By matching and combining the recorded values ​​of the same key value in the data set, the optimal matching combination is preferred according to the combined distance standard deviation, the data matching problem when the non-error distance deviation is large in the prior art is solved, and efficient data matching effect is achieved.

CN114048247BActive Publication Date: 2025-05-09CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111354785.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-16
Publication Date
2025-05-09
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

The prior art cannot meet the data matching requirements when the non-error distance deviation is large, which may lead to the problem of data matching association misalignment.

Method used

By matching any two recorded values ​​of the same key value in the two data sets to be matched, all possible matching combinations are found, and then the optimal matching combination is selected based on the combined distance standard deviation, reducing or eliminating the effects of non-error distance deviation.

Benefits of technology

The data matching between two data sets with large non-error distance deviation is achieved, avoiding omissions or incorrect correlation matching, and meeting the data matching needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114048247B_ABST
    Figure CN114048247B_ABST
Patent Text Reader

Abstract

The present invention provides a data matching method, device, computer equipment and readable storage medium, the method comprising: obtaining key values ​​for association and record values ​​for matching for two data sets to be matched; matching any two record values ​​of the same key value from the two data sets according to their mutual distances; combining all successfully matched record value pairs to obtain all possible matching combinations of the same key value; sequentially obtaining all possible matching combinations of all the same key values ​​in the two data sets to be matched; obtaining the combined distance standard deviation for each possible matching combination of the same key value according to the distance between all its record value pairs, and determining the optimal matching combination of the same key value according to the combined distance standard deviation; sequentially obtaining the optimal matching combinations of all the same key values ​​of the two data sets to be matched to obtain data matching results. The present invention can also meet the demand for accurate data matching when the non-error distance deviation is large.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data matching technology, and in particular to a data matching method, device, computer equipment and readable storage medium. Background Art

[0002] With the rise of cloud computing and big data industries, all walks of life have paid more attention to the value of data. When applying cross-domain data, it is necessary to associate the data generated by different data generation systems to achieve data matching. Different data generation systems have error distance deviation and non-error distance deviation for the same data. The non-error distance deviation is the data misalignment caused by the time deviation and measurement deviation of different data generation systems.

[0003] The existing general technology usually uses judgment conditions such as keyword inclusion and sameness to directly compare the values ​​to achieve association matching, which cannot meet the matching requirements between data with non-error distance deviation. The applicant has proposed a data matching method and device (CN108920601A). However, after verification, it was found that when the non-error distance deviation is greater than the error distance deviation, this matching algorithm may fail, which will lead to the problem of data association misalignment. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a data matching method, device, computer equipment and readable storage medium in response to the above-mentioned deficiencies in the prior art, so as to solve the problem that the prior art cannot meet the data matching requirements or may cause data matching association misalignment when the non-error distance deviation is large.

[0005] In a first aspect, the present invention provides a data matching method, comprising:

[0006] Obtain the key value used for association and the record value used for matching for the two data sets to be matched;

[0007] Match any two record values ​​of the same key value from two data sets according to their mutual distance;

[0008] Combine all successfully matched record value pairs to obtain all possible matching combinations of the same key value;

[0009] Sequentially obtain all possible matching combinations of all the same key values ​​in the two data sets to be matched;

[0010] For each possible matching combination of the same key value, obtain the combined distance standard deviation based on the distance between all its record value pairs, and determine the optimal matching combination of the same key value based on the combined distance standard deviation;

[0011] The optimal matching combination of all the same key values ​​of the two data sets to be matched is obtained in sequence to obtain the data matching result.

[0012] Preferably, after obtaining the key value for association and the record value for matching for the two data sets to be matched, the method further comprises:

[0013] All record values ​​of the two data sets to be matched are grouped according to the key value, and all record values ​​in each group are sorted.

[0014] Preferably, after grouping all record values ​​of the two data sets to be matched according to the key values, the method further comprises:

[0015] Select one of the data sets and calculate its average sampling point spacing according to the following formula:

[0016]

[0017] Among them, D msd is the average sampling point spacing of the data set, G is the number of groups in the data set, m i is the number of record values ​​of the i-th group of the data set, is the xth record value of the ith group of the data set, is the x-1th record value of the i-th group of the data set;

[0018] The average sampling point spacing of the data set is compared with the average sampling point spacing preset by the user. If the average sampling point spacing of the data set is greater than or equal to the average sampling point spacing preset by the user, it is determined that the two data sets to be matched can be matched through the data matching method, and the subsequent steps are continued. Otherwise, it is determined that the two data sets to be matched cannot be matched through the data matching method, and the data matching method is terminated.

[0019] Preferably, the matching of any two record values ​​of the same key value from two data sets according to their mutual distances specifically includes:

[0020] Get all record values ​​of the same key value from two data sets;

[0021] Set up an outer loop in the sorted order of all records with the same key value from one of the data sets;

[0022] Setting an inner loop nested within the outer loop according to the sorted order of all record values ​​of the same key value from another data set;

[0023] When the distance between two record values ​​obtained by the inner loop calculation meets a preset condition, it is determined that the corresponding two record values ​​can be successfully matched.

[0024] Preferably, the preset condition is that the distance between them is less than the standard deviation of the magnified error distance σ resf ;

[0025] Before obtaining all record values ​​of the same key value from two data sets, the method further includes:

[0026] The preset maximum error distance standard deviation σ res According to the magnification factor f mag After magnification, the magnification error distance standard deviation σ is obtained resf .

[0027] Preferably, the matching of any two record values ​​of the same key value from two data sets according to their mutual distances further includes:

[0028] When the distance between the two recorded values ​​calculated in this inner loop is greater than or equal to the standard deviation of the amplified error distance σ resf , and when the index value of the inner loop is greater than the index value of the outer loop, the inner loop ends.

[0029] Preferably, the combining of all successfully matched record value pairs to obtain all possible matching combinations of the same key value is specifically as follows:

[0030] All successfully matched record value pairs are combined according to the principle of the largest number of record value pairs in each matching combination without duplication, to obtain all possible matching combinations of the same key value.

[0031] Preferably, for each possible matching combination of the same key value, the standard deviation of the combined distance is obtained according to the distance between all the record value pairs thereof, which is specifically obtained according to the following formula:

[0032]

[0033] Among them, μ is the mean distance between all recorded value pairs, D k is the distance between the kth pair of record values, K is the number of all record value pairs, and σ is the standard deviation of the combined distance.

[0034] Preferably, determining the optimal matching combination of the same key value according to the combination distance standard deviation specifically includes:

[0035] When there are multiple combined distance standard deviations σ that are less than the preset maximum error distance standard deviation σ resWhen , select the matching combination with more record value pairs and other matching combinations with less than or equal to the number of record pairs to find the intersection and complement, and calculate the combined distance standard deviation of the complement. If the combined distance standard deviation of the complement is also less than the preset maximum error distance standard deviation σ res When , the matching combination with more record value pairs is selected as the optimal matching combination, otherwise the matching combination with fewer record value pairs is selected as the optimal matching combination.

[0036] In a second aspect, the present invention provides a data matching device, comprising:

[0037] An acquisition module, used for acquiring a key value for association and a record value for matching from two data sets to be matched;

[0038] A matching module, connected to the acquisition module, for matching any two record values ​​of the same key value from two data sets according to the distance between them;

[0039] A combination module, connected to the matching module, for combining all successfully matched record value pairs to obtain all possible matching combinations of the same key value;

[0040] A cyclic acquisition module, connected to the combination module and the matching module, for sequentially acquiring all possible matching combinations of all the same key values ​​in the two data sets to be matched;

[0041] A preferred module, connected to the loop obtaining module, is used to obtain a combination distance standard deviation for each possible matching combination of the same key value according to the distances between all its record value pairs, and determine the optimal matching combination of the same key value according to the combination distance standard deviation;

[0042] The optimization module is circulated and connected with the optimization module to sequentially obtain the optimal matching combination of all the same key values ​​of the two data sets to be matched to obtain the data matching result.

[0043] In a third aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the data matching method as described above.

[0044] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the data matching method as described above is implemented.

[0045] The present invention provides a data matching method, device, computer equipment and readable storage medium, which matches any two record values ​​of the same key value in two data sets to be matched to find all possible matching combinations, and then selects the best matching combination from all possible matching combinations of the same key value by comparing the combined distance standard deviation, reduces or eliminates the influence of the non-error distance deviation by combining the distance standard deviation, and performs the above matching, combining and selecting on all record values ​​of the same key value in the two data sets to be matched, thereby obtaining all the best matching methods between all record values ​​of the two data sets to be matched, avoiding omissions or erroneous associated matching, and the above method can meet the data matching between two data sets with large non-error distance deviations. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is an example diagram of application scenarios of the present invention;

[0047] Figure 2 is a flow chart of a data matching method according to embodiment 1 of the present invention;

[0048] Figure 3 is a specific flow chart of step S2 of embodiment 1 of the present invention;

[0049] Figure 4 is a structural schematic diagram of a data matching device according to embodiment 2 of the present invention;

[0050] Figure 5 It is a structural diagram of a computer device according to Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0051] In order to enable those skilled in the art to better understand the technical solution of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0052] It should be understood that the specific embodiments and drawings described herein are only used to explain the present invention rather than to limit the present invention.

[0053] It can be understood that, in the absence of conflict, the various embodiments of the present invention and the various features in the embodiments can be combined with each other.

[0054] It can be understood that, for the convenience of description, the drawings of the present invention only show the parts related to the present invention, while the parts irrelevant to the present invention are not shown in the drawings.

[0055] It can be understood that each unit and module involved in the embodiments of the present invention may correspond to only one physical structure, or may be composed of multiple physical structures, or multiple units and modules may be integrated into one physical structure.

[0056] It can be understood that, in the absence of conflict, the functions and steps marked in the flowcharts and block diagrams of the present invention may occur in an order different from that marked in the drawings.

[0057] It is understood that the flowcharts and block diagrams of the present invention illustrate the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the various embodiments of the present invention. Each box in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified functions. Moreover, each box or combination of boxes in the block diagram and flowchart may be implemented by a hardware-based system that implements the specified functions, or may be implemented by a combination of hardware and computer instructions.

[0058] It can be understood that the units and modules involved in the embodiments of the present invention can be implemented by software or hardware. For example, the units and modules can be located in a processor.

[0059] In order to facilitate the understanding of the present invention, the generation of non-error distance deviation is first explained. At present, many applications in the ICT (Information And Communication Technology) industry, such as the association of data records between the billing domain and the operation and maintenance domain, the association of detailed signaling records between the wireless network and the core network, often use cross-domain data.

[0060] like Figure 1 As shown, it is an example of one of the application scenarios of the present invention, which shows the non-error distance deviation between RRC (Radio Resource Control) signaling and core network signaling. Specifically, for a certain RRC signaling, after the core network receives and processes the RRC signaling, it will generate a corresponding core network signaling. The RRC signaling is tracked by the TRACE tracking system on the wireless side, and the core network signaling is tracked by the DPI (Deep Package Inspection) system on the core network side. There is a transmission delay from the wireless network to the core network. It takes a certain amount of time for the core network to process the RRC signaling and generate the core network signaling, t trans T represents the transmission delay of signaling from the wireless network to the core network. process The time taken by the core network to process RRC signaling and generate core network signaling. At the same time, the clock settings of the TRACE tracking system and the DPI system may not be aligned, resulting in time deviation t d , they together constitute the non-error time deviation. When this non-error time deviation is large and has not been located, it will cause a certain RRC signaling to incorrectly match the core network signaling corresponding to other RRC signaling.

[0061] Embodiment 1:

[0062] like Figure 2 As shown, Embodiment 1 of the present invention provides a data matching method, including:

[0063] Step S1: Obtain key values ​​for association and record values ​​for matching for two data sets to be matched.

[0064] For example, this embodiment 1 can be used to perform data matching on user call records in the billing domain and the operation and maintenance domain. The user call records from the billing domain are recorded as data set A, and the user call records from the operation and maintenance domain are recorded as data set B. First, key values ​​for association and record values ​​for matching are selected for data set A and data set B. For example, if this embodiment 1 is to match the call record data of the same calling user in data set A and data set B, the calling user phone number in the user call record can be selected as the key value for association according to the matching requirements, and the start time of the call connection in the call record can be selected as the record value for matching.

[0065] In this embodiment, after obtaining the key value for association and the record value for matching for the two data sets to be matched in step S1, the method further includes:

[0066] All record values ​​of the two data sets to be matched are grouped according to the key value, and all record values ​​in each group are sorted.

[0067] Specifically, in this embodiment 1, for the convenience of description and calculation, the key value and the record value are parameterized. For data set A, its user call records contain call records of M different calling user phone numbers, where phone number 1 contains m i In this call record, the key values ​​of M different calling user phone numbers are marked as KA i (1≤i≤M), the phone number 1 contains m i The call records are marked as For data set B, its user call records contain call records of N different calling user phone numbers, where phone number 1 contains n j In this call record, the key values ​​of N different calling user phone numbers are marked as KB j (1≤j≤N), the phone number 1 contains n j The call records are marked as (1≤y≤n j ).

[0068] To make it easier to understand, let's assume a simple data example.

[0069]

[0070] The meaning of the above table is that after obtaining the key value of data set A, we get KA1-KA M There are M different key values ​​in total. Similarly, for data set B, we get KB1-KB N There are N different key values, among which KA1 = KB1 = telephone number 1 (the subscript value can be different, by i and KB j By performing a circular comparison, all the same key values ​​in data set A and data set B can be found. For simplicity, the subscript values ​​are all 1 here), and in data set A, phone number 1 has three call records, and the recorded call start times are 3:00, 3:01, and 3:05, respectively. They are assigned to the group marked with KA1 in chronological order. In data set B, phone number 1 also has three call records (the number of call records may be unequal due to omissions, etc., and it is assumed to be equal here), and the recorded call start times are 3:01, 3:02, and 3:06, respectively. They are assigned to the group marked with KB1 in chronological order.

[0071] In this embodiment, after grouping all record values ​​of the two data sets to be matched according to the key values, the method further includes:

[0072] Select one of the data sets and calculate its average sampling point spacing according to the following formula:

[0073]

[0074] Among them, D msd is the average sampling point spacing of the data set, G is the number of groups in the data set, m i is the number of record values ​​of the i-th group of the data set, is the xth record value of the ith group of the data set, is the x-1th record value of the i-th group of the data set;

[0075] The average sampling point spacing of the data set is compared with the average sampling point spacing preset by the user. If the average sampling point spacing of the data set is greater than or equal to the average sampling point spacing preset by the user, it is determined that the two data sets to be matched can be matched through the data matching method, and the subsequent steps are continued. Otherwise, it is determined that the two data sets to be matched cannot be matched through the data matching method, and the data matching method is terminated.

[0076] Step S2: Match any two record values ​​of the same key value from two data sets according to the distance between them.

[0077] In this embodiment, if Figure 3 As shown, step S2 matches any two record values ​​of the same key value from two data sets according to their mutual distances, specifically including:

[0078] Step S21: Obtain all record values ​​of the same key value from two data sets respectively;

[0079] Step S22: setting an outer loop according to the sorted order of all record values ​​of the same key value from one of the data sets;

[0080] Step S23: setting an inner loop nested in the outer loop according to the sorted order of all record values ​​of the same key value from another data set;

[0081] Step S24: when the distance between two record values ​​obtained by the inner loop calculation meets the preset condition, it is determined that the corresponding two record values ​​can be successfully matched.

[0082] Specifically, in the present embodiment 1, step S21 obtains all record values ​​of the same key value from two data sets, that is, obtains all record values ​​of two groups with equal key values. In the above simple data example, according to KA1=KB1 (it can also be KA2=KB2, etc., only one set of examples is given here) Set the nested loop calculation according to the sorted order, which can be controlled by setting (1≤x≤3) The outer loop of the control is set (1≤y≤3) The inner loop of , thereby realizing the distance calculation between any two record values, taking the above example as an example, the distance between any two record values ​​is calculated through the loop. D 12 =2,D 13 =6, D 21 =0,D 22 =1,D 23 =5,D 31 =4,D 32 =3,D 33 =1 (in minutes).

[0083] In this embodiment, the preset condition is that the distance between them is less than the standard deviation of the magnification error distance σ resf ;

[0084] Before obtaining all record values ​​of the same key value from two data sets in step S21, the method further includes:

[0085] The preset maximum error distance standard deviation σ res According to the magnification factor f magAfter magnification, the standard deviation of the magnified error distance σ is obtained resf .

[0086] Specifically, in this embodiment 1, for example, σ is set res =1.5, f mag =3, then σ resf =4.5, then in the above example, the record values ​​that can be successfully matched include:

[0087] Specifically, in this embodiment 1, Figure 3 As shown, step S2 matches any two record values ​​of the same key value from two data sets according to their mutual distances, and specifically includes:

[0088] Step S25: When the distance between two record values ​​calculated by the inner loop does not meet the preset condition, and the index value of the inner loop is greater than the index value of the outer loop, or the index value of the inner loop reaches the maximum preset value, the inner loop is terminated;

[0089] Step S26: When the index value of the outer loop reaches the maximum preset value, the outer loop is terminated to obtain all successfully matched record value pairs.

[0090] In this embodiment, when the distance between the two recorded values ​​calculated by the inner loop does not meet the preset conditions, specifically: when the distance between the two recorded values ​​calculated by the inner loop is greater than or equal to the standard deviation of the magnified error distance σ resf .

[0091] Specifically, in this embodiment 1, when the index value of the inner loop reaches the maximum preset value, the inner loop ends, and then the next inner loop starts. When the index value of the outer loop reaches the maximum preset value, the outer loop ends, so that the entire nested loop calculation ends. This is the characteristic of the loop calculation itself. This embodiment 1 is mainly for the sorted record values. When the current record value pair does not meet the preset conditions, the inner loop of this round can be ended in advance to reduce the amount of calculation. For the above example, if there are still record values: and When x=1, calculate to D 13 =6>σ resf =4.5, and at this time y=3>x=1, then there is no need to calculate D 14 , directly end this round of inner loop, set x=2 to enter the next round of inner loop. This is because the record values ​​are arranged in order (in ascending order in this example). When the previous pairing can no longer meet the requirements, the next pairing will inevitably not meet the requirements.

[0092] Step S3: Combine all successfully matched record value pairs to obtain all possible matching combinations of the same key value.

[0093] In this embodiment, step S3 combines all successfully matched record value pairs to obtain all possible matching combinations of the same key value, specifically:

[0094] All successfully matched record value pairs are combined according to the principle of the largest number of record value pairs in each matching combination without duplication, to obtain all possible matching combinations of the same key value.

[0095] Specifically, in this embodiment 1, for the above example, for the same key value KA1=KB1, according to the principle of the maximum number of record value pairs in each matching combination without duplication, all possible matching combinations are obtained as follows:

[0096] Step S4: sequentially obtain all possible matching combinations of all the same key values ​​in the two data sets to be matched.

[0097] Specifically, in this embodiment 1, for the above example, if there are other groups with the same key value in data set A and data set B, steps S2 and S3 are repeated in a loop to obtain all possible matching combinations of each group with the same key value.

[0098] Step S5: for each possible matching combination of the same key value, obtain the combination distance standard deviation according to the distances between all its record value pairs, and determine the optimal matching combination of the same key value according to the combination distance standard deviation.

[0099] In this embodiment, for each possible matching combination of the same key value in step S5, the standard deviation of the combined distance is obtained according to the distance between all the record value pairs, which is specifically obtained according to the following formula:

[0100]

[0101] Among them, μ is the mean distance between all record value pairs, Dk is the distance between the kth pair of record values, K is the number of all record value pairs, and σ is the standard deviation of the combined distance.

[0102] Specifically, since the error distance deviation is volatile and the non-error distance deviation is stable, in this embodiment 1, the combined distance standard deviation σ of all record value pairs for each possible matching combination of the same key value is calculated, and the calculation of the combined distance standard deviation σ subtracts the mean μ of the distances between all record value pairs in the matching combination, thereby reducing the impact of the non-error distance deviation on the matching result.

[0103] In this embodiment, determining the optimal matching combination of the same key value according to the combination distance standard deviation in step S5 specifically includes:

[0104] When there are multiple combined distance standard deviations σ that are less than the preset maximum error distance standard deviation σ res When , select the matching combination with more record value pairs and other matching combinations with less than or equal to the number of record pairs to find the intersection and complement, and calculate the combined distance standard deviation of the complement. If the combined distance standard deviation of the complement is also less than the preset maximum error distance standard deviation σ res When , the matching combination with more record value pairs is selected as the optimal matching combination, otherwise the matching combination with fewer record value pairs is selected as the optimal matching combination.

[0105] Specifically, in this embodiment 1, the combined distance standard deviation of the complement set is calculated for the Z pairs of record values ​​in the obtained complement set according to the following formula:

[0106]

[0107] The formula is the same as that for each possible matching combination, the standard deviation of the combined distance is obtained based on the distance between all the record value pairs. However, only the deviation of the complement part is calculated here, and the complement part is determined to be within the preset maximum error distance standard deviation σ. res If , it means that the record value pairs in the complement set also meet the error standard, so the matching combination with more pairs is a better choice.

[0108] For the above example, C1: σ = 0, C2: σ = 1, C3: σ = 2 / 3, C5: σ = 1, all of which are less than σ res =1.5. In this example, since C1 has the smallest σ and the largest number of logarithms, it can actually be determined as the optimal matching result. However, for more complex data records, further calculation is required. For example, if only C3 and C5 meet the requirements in this example, how to determine which one is better, C3 or C5, then it is necessary to compare C3 and C5. The intersection of C3 and C5 is The intersection of C3 with C5 is Then C3 is a better matching combination. At this time, being able to calculate the combined distance standard deviation of the complement means that the number of record value pairs in the complement is greater than 1. If it is 1, you can select any record value pair from the intersection and add it as the complement to calculate the combined distance standard deviation.

[0109] Step S6: sequentially obtain the optimal matching combination of all the same key values ​​of the two data sets to be matched, and obtain the data matching result.

[0110] Specifically, in this embodiment 1, for the above example, if there are other groups with the same key value in data set A and data set B, step S5 is repeated in a loop to obtain the optimal matching combination of each group with the same key value, thereby obtaining a data matching result.

[0111] Embodiment 2:

[0112] like Figure 4 As shown, Embodiment 2 of the present invention provides a data matching device, including:

[0113] Acquisition module 1 is used to obtain key values ​​for association and record values ​​for matching for two data sets to be matched:

[0114] A matching module 2, connected to the acquisition module 1, is used to match any two record values ​​of the same key value from two data sets according to the distance between them;

[0115] A combination module 3, connected to the matching module 2, is used to combine all successfully matched record value pairs to obtain all possible matching combinations of the same key value;

[0116] A loop obtaining module 4, connected to the combination module 3 and the matching module 2, for sequentially obtaining all possible matching combinations of all the same key values ​​in the two data sets to be matched;

[0117] The preferred module 5 is connected to the loop obtaining module 4, and is used to obtain the combination distance standard deviation for each possible matching combination of the same key value according to the distance between all its record value pairs, and determine the optimal matching combination of the same key value according to the combination distance standard deviation;

[0118] The optimization module 5 is circulated and connected to the optimization module 5 to sequentially obtain the optimal matching combination of all the same key values ​​of the two data sets to be matched to obtain the data matching result.

[0119] Optionally, the data matching device further includes:

[0120] The sorting module is used to group all the record values ​​of the two data sets to be matched according to the key value, and sort all the record values ​​in each group.

[0121] Optionally, the data matching device further includes:

[0122] The judgment module is used to:

[0123] Select one of the data sets and calculate its average sampling point spacing according to the following formula:

[0124]

[0125] Among them, D msd is the average sampling point spacing of the data set, G is the number of groups of the data set, mi is the number of record values ​​of the i-th group of the data set, is the xth record value of the ith group of the data set, is the x-1th record value of the i-th group of the data set;

[0126] The average sampling point spacing of the data set is compared with the average sampling point spacing preset by the user. If the average sampling point spacing of the data set is greater than or equal to the average sampling point spacing preset by the user, it is determined that the two data sets to be matched can be matched through the data matching method, and the subsequent steps are continued. Otherwise, it is determined that the two data sets to be matched cannot be matched through the data matching method, and the data matching method is terminated.

[0127] Optionally, the matching module 2 specifically includes:

[0128] An acquisition unit, used for acquiring all record values ​​of the same key value from two data sets respectively;

[0129] An outer loop unit, used for setting an outer loop according to the sorted order of all record values ​​of the same key value from one of the data sets;

[0130] An inner loop unit, used for setting an inner loop nested in the outer loop according to the sorted order of all record values ​​of the same key value from another data set;

[0131] The determination unit is used to determine that the corresponding two record values ​​can be successfully matched when the distance between the two record values ​​obtained by the inner loop calculation meets a preset condition.

[0132] Optionally, the preset condition is that the distance between them is less than the standard deviation of the magnified error distance σ resf ;

[0133] The data matching device also includes:

[0134] The preset module is used to set the preset maximum error distance standard deviation σ res According to the magnification factor f mag After magnification, the magnification error distance standard deviation σ is obtained resf .

[0135] Optionally, the matching module 2 specifically further includes:

[0136] The inner loop ending unit is used to end the inner loop when the distance between the two recorded values ​​calculated by the inner loop does not meet the preset conditions, and the index value of the inner loop is greater than the index value of the outer loop, or the index value of the inner loop reaches the maximum preset value; the inner loop calculation when the distance between the two recorded values ​​does not meet the preset conditions is specifically: when the distance between the two recorded values ​​calculated by the inner loop is greater than or equal to the standard deviation of the amplified error distance σ resf ;

[0137] The outer loop ending unit is used to end the outer loop when the index value of the outer loop reaches a maximum preset value, and obtain all successfully matched record value pairs.

[0138] Optionally, the combination module 3 is specifically used for:

[0139] All successfully matched record value pairs are combined according to the principle of the largest number of record value pairs in each matching combination without duplication, to obtain all possible matching combinations of the same key value.

[0140] Optionally, the preferred module 5 includes:

[0141] The standard deviation calculation unit is used to obtain the standard deviation of the combined distance for each possible matching combination of the same key value according to the distance between all the record value pairs, which is specifically obtained according to the following formula:

[0142]

[0143] Among them, μ is the mean distance between all recorded value pairs, D k is the distance between the kth pair of record values, K is the number of all record value pairs, and σ is the standard deviation of the combined distance.

[0144] Optionally, the preferred module 5 further includes:

[0145] Determine the optimal unit, which is used to determine the optimal matching combination of the same key value based on the standard deviation of the combination distance, specifically used for:

[0146] When there are multiple combined distance standard deviations σ that are less than the preset maximum error distance standard deviation σ res When , select the matching combination with more record value pairs and other matching combinations with less than or equal to the number of record pairs to find the intersection and complement, and calculate the combined distance standard deviation of the complement. If the combined distance standard deviation of the complement is also less than the preset maximum error distance standard deviation σ res When , the matching combination with more record value pairs is selected as the optimal matching combination, otherwise the matching combination with fewer record value pairs is selected as the optimal matching combination.

[0147] Embodiment 3:

[0148] like Figure 5 As shown, embodiment 3 of the present invention provides a computer device, which includes a memory 10 and a processor 20. The memory 10 stores a computer program. When the processor 20 runs the computer program stored in the memory 10, the processor 20 executes the data matching method described in embodiment 1.

[0149] The memory 10 is connected to the processor 20. The memory 10 may be a flash memory, a read-only memory or other memory. The processor 20 may be a central processing unit or a single-chip microcomputer.

[0150] Embodiment 4:

[0151] Embodiment 4 of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the data matching method as described in Embodiment 1 is implemented.

[0152] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules or other data). Computer-readable storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable read only memory), flash memory or other memory technology, CD-ROM (Compact Disc Read-Only Memory), digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer.

[0153] Embodiments 1-4 of the present invention provide a data matching method, apparatus, computer equipment and readable storage medium, which matches any two record values ​​of the same key value in two data sets to be matched to find all possible matching combinations, and then selects the best matching combination from all possible matching combinations of the same key value by comparing the combined distance standard deviation, reduces or eliminates the influence of the non-error distance deviation by combining the distance standard deviation, and performs the above matching, combining and selecting on all record values ​​of the same key value in the two data sets to be matched, thereby obtaining all the best matching methods between all record values ​​of the two data sets to be matched, avoiding omissions or erroneous associated matches, and the above method can meet the data matching between two data sets with large non-error distance deviations.

[0154] It is to be understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of the present invention, but the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered to be within the scope of protection of the present invention.

Claims

1. A data matching method, characterized in that: include: Obtain the key value used for association and the record value used for matching for the two data sets to be matched; Group all the record values ​​of the two data sets to be matched according to the key value, and sort all the record values ​​in each group; Select one of the grouped and sorted data sets and calculate its average sampling point spacing according to the following formula: Among them, D msd is the average sampling point spacing of the data set, G is the number of groups in the data set, m i is the number of record values ​​of the i-th group of the data set, is the xth record value of the ith group of the data set, is the x-1th record value of the ith group of the data set, Compare the average sampling point spacing of the data set with the average sampling point spacing preset by the user, if the average sampling point spacing of the data set is greater than or equal to the average sampling point spacing preset by the user, continue to execute the subsequent steps, otherwise, do not execute the subsequent steps; Any two record values ​​of the same key value from two data sets are matched according to the distance between them. When the distance between them meets the preset condition, it is determined that the corresponding two record values ​​can be successfully matched. The preset condition is that the distance between them is less than the standard deviation of the magnified error distance σ resf ; Combine all successfully matched record value pairs to obtain all possible matching combinations of the same key value; Sequentially obtain all possible matching combinations of all the same key values ​​in the two data sets to be matched; For each possible matching combination of the same key value, obtain the combined distance standard deviation based on the distance between all its record value pairs, and determine the optimal matching combination of the same key value based on the combined distance standard deviation; The optimal matching combination of all the same key values ​​of the two data sets to be matched is obtained in sequence to obtain the data matching result.

2. The data matching method according to claim 1, characterized in that: The matching of any two record values ​​of the same key value from two data sets according to the distance between them specifically includes: Get all record values ​​of the same key value from two data sets; Set up an outer loop in the sorted order of all records with the same key value from one of the data sets; Setting an inner loop nested within the outer loop according to the sorted order of all record values ​​of the same key value from another data set; When the distance between two record values ​​obtained by the inner loop calculation meets a preset condition, it is determined that the corresponding two record values ​​can be successfully matched.

3. The data matching method according to claim 2, characterized in that: Before obtaining all record values ​​of the same key value from two data sets, the method further includes: The preset maximum error distance standard deviation σ res According to the magnification factor f mag After magnification, the magnification error distance standard deviation σ is obtained resf .

4. The data matching method according to claim 3, characterized in that: The matching of any two record values ​​of the same key value from two data sets according to the distance between them specifically includes: When the distance between two recorded values ​​calculated by the inner loop is greater than or equal to the standard deviation of the magnified error distance σ resf , and when the index value of the inner loop is greater than the index value of the outer loop, the inner loop ends.

5. The data matching method according to claim 4, characterized in that: The combination of all successfully matched record value pairs to obtain all possible matching combinations of the same key value is specifically: All successfully matched record value pairs are combined according to the principle of the largest number of record value pairs in each matching combination without duplication, to obtain all possible matching combinations of the same key value.

6. The data matching method according to claim 5, characterized in that: For each possible matching combination of the same key value, the standard deviation of the combined distance is obtained according to the distance between all the record value pairs, which is specifically obtained according to the following formula: Among them, μ is the mean distance between all recorded value pairs, D k is the distance between the kth pair of record values, K is the number of all record value pairs, and σ is the standard deviation of the combined distance.

7. The data matching method according to claim 6, characterized in that: Determining the optimal matching combination of the same key value according to the combination distance standard deviation specifically includes: When there are multiple combined distance standard deviations σ that are less than the preset maximum error distance standard deviation σ res When , select the matching combination with more record value pairs and other matching combinations with less than or equal to the number of record pairs to find the intersection and complement, and calculate the combined distance standard deviation of the complement. If the combined distance standard deviation of the complement is also less than the preset maximum error distance standard deviation σ res When , the matching combination with more record value pairs is selected as the optimal matching combination, otherwise the matching combination with fewer record value pairs is selected as the optimal matching combination.

8. A data matching device, characterized in that: include: An acquisition module, used for acquiring a key value for association and a record value for matching from two data sets to be matched; A sorting module, connected to the acquisition module, for grouping all record values ​​of the two data sets to be matched according to the key values, and sorting all record values ​​in each group; The judgment module is connected to the sorting module and is used to select one of the grouped and sorted data sets and calculate its average sampling point spacing according to the following formula: Among them, D msd is the average sampling point spacing of the data set, G is the number of groups in the data set, m i is the number of record values ​​of the i-th group of the data set, is the xth record value of the ith group of the data set, is the x-1th record value of the ith group of the data set, Compare the average sampling point spacing of the data set with the average sampling point spacing preset by the user, if the average sampling point spacing of the data set is greater than or equal to the average sampling point spacing preset by the user, continue to execute the subsequent steps, otherwise, do not execute the subsequent steps; A matching module is connected to the judgment module and is used to match any two record values ​​of the same key value from two data sets according to the distance between them. When the distance between them meets the preset condition, it is determined that the corresponding two record values ​​can be successfully matched. The preset condition is that the distance between them is less than the standard deviation of the magnified error distance σ resf ; A combination module, connected to the matching module, for combining all successfully matched record value pairs to obtain all possible matching combinations of the same key value; A cyclic acquisition module, connected to the combination module and the matching module, for sequentially acquiring all possible matching combinations of all the same key values ​​in the two data sets to be matched; A preferred module, connected to the loop obtaining module, is used to obtain a combination distance standard deviation for each possible matching combination of the same key value according to the distances between all its record value pairs, and determine the optimal matching combination of the same key value according to the combination distance standard deviation; The optimization module is circulated and connected with the optimization module to sequentially obtain the optimal matching combination of all the same key values ​​of the two data sets to be matched to obtain the data matching result.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor runs the computer program stored in the memory, the processor executes the data matching method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the data matching method as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Data matching method and device

    CN108920601A

  • Intelligent image matching method and device and computer readable storage medium

    CN110633733A

  • Machine learnt match rules

    US20190236460A1