Call Detail Record Duplicate Detection Method, Device, and Computer-Readable Storage Medium
By generating and optimizing the combination of key fields for plagiarism checking, the problems of low efficiency and insufficient plagiarism checking in the prior art are solved, and efficient and accurate plagiarism checking and whole price processing are achieved.
Patent Information
- Application Number
- CN202111179081.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-10-08
AI Technical Summary
The prior art cannot automatically select the optimal key field combination during the diversion checking process of the call-book, resulting in low diversion checking efficiency, and manual-designed diversion checking fields are prone to redundancy or lack, affecting the diversion checking accuracy and wholesale price accuracy.
By obtaining the key fields and preset fields associated with the repetition list corresponding to the current cycle, an initial key field combination is generated, and the risk coefficient of each initial key field combination is determined based on the preset key field check risk coefficient function model, the target key field combination is selected, and the repetition list check is performed.
It realizes the automatic selection of the optimal key field combination based on the actual call order situation, improves the efficiency of call order checking, and ensures the accuracy of the checking of pounds and the accuracy of the whole price.
Smart Images

Figure CN115967768B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and particularly to a call record duplicate checking method, a call record duplicate checking device, and a computer-readable storage medium. Background Art
[0002] The process of call record warehousing for network operator billing and settlement includes call record collection, decoding, analysis, rating, and warehousing. Due to repeated collection caused by network problems and other reasons, call records will be checked for duplicates and then rated after the analysis step.
[0003] In related technologies, when it is necessary to check call records for duplicates, manual selection of duplicate checking indexes from a large number of call record fields is usually adopted, and the optimal keyword field combination cannot be automatically selected according to the actual call record situation, resulting in low efficiency of call record duplicate checking.
[0004] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention
[0005] The main purpose of the present invention is to provide a call record duplicate checking method, a call record duplicate checking device, and a computer-readable storage medium, aiming to achieve the effect of improving the effectiveness of the operator's evaluation of the user experience.
[0006] To achieve the above purpose, the present invention provides a call record duplicate checking method, which includes the following steps:
[0007] Obtain keyword fields associated with call records to be checked for duplicates corresponding to the current period and a preset field combination;
[0008] Generate an initial keyword field combination based on the preset field combination and the keyword fields;
[0009] Determine the risk coefficient corresponding to each initial keyword field combination according to a preset keyword field duplicate checking risk coefficient function model;
[0010] Select a target keyword field combination from the initial keyword field combinations according to the risk coefficient;
[0011] Perform call record duplicate checking on the call records to be checked for duplicates corresponding to the current period based on the target keyword field combination.
[0012] Optionally, before the step of determining the risk coefficient corresponding to each initial keyword field combination according to a preset keyword field duplicate checking risk coefficient function model, the method further includes:
[0013] Obtain the first parameter estimation result and the first system observation information corresponding to the previous period;
[0014] Calculate the Kalman gain based on the estimation result of the first parameter and the first system observation information;
[0015] Determine the optimization parameter according to the Kalman gain and the system measurement parameter corresponding to the current period;
[0016] Optimize the parameters of the preset keyword field duplicate check risk coefficient function model according to the optimization parameter.
[0017] Optionally, after the step of calculating the Kalman gain according to the estimation result of the first parameter and the first system observation information, the method further includes:
[0018] Determine the estimation result of the second parameter and the second system observation information corresponding to the next period according to the Kalman gain and the system measurement parameter corresponding to the current period, and save the estimation result of the second parameter and the second system observation information.
[0019] Optionally, the step of performing duplicate check on the call records to be checked corresponding to the current period based on the target keyword field combination includes:
[0020] Determine the first call record information digest corresponding to the prior call record and the second call record information digest corresponding to the call record to be checked based on the target keyword field combination;
[0021] Determine whether the call record to be checked is a duplicate call record according to the first call record information digest and the second call record information digest.
[0022] Optionally, the step of determining the first call record information digest corresponding to the prior call record and the second call record information digest corresponding to the call record to be checked based on the target keyword field combination includes:
[0023] Determine the index information corresponding to the prior call record according to the target keyword field combination;
[0024] Use the index information as the input parameter of the Secure Hash Algorithm 1 (SHA1), calculate the string corresponding to the index information through the SHA1 algorithm, and use the string as the first call record information digest; and
[0025] Calculate the second call record information digest corresponding to the call record to be checked through the SHA1 algorithm based on the target keyword field combination.
[0026] Optionally, after the step of performing duplicate check on the call records to be checked corresponding to the current period based on the target keyword field combination, the method further includes:
[0027] Determine the normal call records according to the call record duplicate check result, and perform rating processing on the normal call records;
[0028] Obtain the pricing amount of the normal call record after pricing processing;
[0029] Perform call record duplicate checking on the normal call record after pricing processing according to the pricing amount and the target keyword field combination to determine whether there are duplicate priced call records.
[0030] Optionally, when performing pricing processing on the normal call record, a token bucket mechanism is used for pricing call record processing.
[0031] In addition, to achieve the above object, the present invention also provides a call record duplicate checking device, which includes: a memory, a processor, and a call record duplicate checking program stored on the memory and executable on the processor. When the call record duplicate checking program is executed by the processor, the steps of the above-mentioned call record duplicate checking method are implemented.
[0032] In addition, to achieve the above object, the present invention also provides a call record duplicate checking device, which includes:
[0033] An acquisition module, configured to acquire keyword fields associated with the call record to be checked for duplicates corresponding to the current period and a preset field combination;
[0034] A generation module, configured to generate an initial keyword field combination based on the preset field combination and the keyword fields;
[0035] A calculation module, configured to determine a risk coefficient corresponding to each initial keyword field combination according to a preset keyword field duplicate checking risk coefficient function model;
[0036] A selection module, configured to select a target keyword field combination from the initial keyword field combinations according to the risk coefficient;
[0037] A duplicate checking module, configured to perform call record duplicate checking on the call record to be checked for duplicates corresponding to the current period based on the target keyword field combination.
[0038] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a call record duplicate checking program is stored. When the call record duplicate checking program is executed by a processor, the steps of the above-mentioned call record duplicate checking method are implemented.
[0039] A call record duplicate checking method, a call record duplicate checking device, and a computer-readable storage medium provided by an embodiment of the present invention first obtain keyword fields associated with call records to be checked corresponding to the current period and a preset field combination, then generate an initial keyword field combination based on the preset field combination and the keyword fields, and then determine a risk coefficient corresponding to each initial keyword field combination according to a preset keyword field duplicate checking risk coefficient function model, select a target keyword field combination from the initial keyword field combinations according to the risk coefficient, and finally perform call record duplicate checking on the call records to be checked corresponding to the current period based on the target keyword field combination. Since the optimal keyword field combination can be automatically selected according to the actual call record situation, the effect of improving the call record duplicate checking efficiency is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic diagram of the terminal structure of the hardware operating environment involved in the solution of the embodiment of the present invention;
[0041] Figure 2 is a schematic diagram of the call record processing flow involved in the related technical solution;
[0042] Figure 3 is a schematic flowchart of an embodiment of the call record duplicate checking method of the present invention;
[0043] Figure 4 is a schematic flowchart of an optional implementation scheme in an embodiment of the call record duplicate checking method of the present invention;
[0044] Figure 5 is a modular schematic diagram of the call record duplicate checking device involved in the embodiment of the present invention.
[0045] The implementation, functional characteristics, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] As Figure 1 shown, Figure 1 is a schematic diagram of the terminal structure of the hardware operating environment involved in the solution of the embodiment of the present invention.
[0048] As Figure 1As shown, the control terminal may include: a processor 1001, such as a CPU, a network interface 1003, a memory 1004, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The network interface 1003 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1004 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1004 may also optionally be a storage device independent of the aforementioned processor 1001.
[0049] Those skilled in the art can understand that Figure 1 the terminal structure shown in
[0050] does not constitute a limitation on the terminal, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 1 As shown in
[0051] In Figure 1 the terminal shown, the processor 1001 may be used to call the call list duplicate checking program stored in the memory 1004 and perform the following operations:
[0052] Obtain the keyword fields and preset field combinations associated with the call list to be checked for the current period;
[0053] Generate an initial keyword field combination based on the preset field combination and the keyword fields;
[0054] Determine the risk coefficient corresponding to each initial keyword field combination according to a preset keyword field duplicate checking risk coefficient function model;
[0055] Select a target keyword field combination from the initial keyword field combinations according to the risk coefficient;
[0056] Perform call list duplicate checking on the call list to be checked for the current period based on the target keyword field combination.
[0057] Further, the processor 1001 may call the call list duplicate checking program stored in the memory 1004 and also perform the following operations:
[0058] Obtain the first parameter estimation result and the first system observation information corresponding to the previous period;
[0059] Calculate the Kalman gain according to the first parameter estimation result and the first system observation information;
[0060] Determine an optimization parameter based on the Kalman gain and the system measurement parameters corresponding to the current period;
[0061] Optimize the parameters of the preset keyword field duplicate check risk coefficient function model according to the optimization parameter.
[0062] Further, the processor 1001 may call the call sheet duplicate check program stored in the memory 1004 and further perform the following operations:
[0063] Determine a second parameter estimation result and second system observation information corresponding to the next period based on the Kalman gain and the system measurement parameters corresponding to the current period, and save the second parameter estimation result and the second system observation information.
[0064] Further, the processor 1001 may call the call sheet duplicate check program stored in the memory 1004 and further perform the following operations:
[0065] Determine a first call sheet information digest corresponding to the prior call sheet and a second call sheet information digest corresponding to the call sheet to be checked for duplicates based on the target keyword field combination;
[0066] Determine whether the call sheet to be checked for duplicates is a duplicate call sheet according to the first call sheet information digest and the second call sheet information digest.
[0067] Further, the processor 1001 may call the call sheet duplicate check program stored in the memory 1004 and further perform the following operations:
[0068] Determine index information corresponding to the prior call sheet according to the target keyword field combination;
[0069] Take the index information as an input parameter of the Secure Hash Algorithm SHA1, calculate a string corresponding to the index information through the SHA1 algorithm, and use the string as the first call sheet information digest; and
[0070] Calculate the second call sheet information digest corresponding to the call sheet to be checked for duplicates through the SHA1 algorithm based on the target keyword field combination.
[0071] Further, the processor 1001 may call the call sheet duplicate check program stored in the memory 1004 and further perform the following operations:
[0072] Determine normal call sheets according to the call sheet duplicate check result, and perform rating processing on the normal call sheets;
[0073] Obtain the rating amount of the normal call sheets after rating processing;
[0074] Perform duplicate check on the normal call records after rating according to the rated amount and the combination of the target key fields to determine whether there are duplicate rated call records.
[0075] The existing charging and settlement call record warehousing usually includes: call record collection, decoding, analysis, rating, and warehousing. Due to repeated collection caused by network problems and other reasons, therefore, duplicate check on call records will be performed after the analysis link and then rating. The main process is as Figure 2 shown. Among them, the call record collection and decoding module is used to collect the original call records generated by network elements of different manufacturers and convert them according to established rules to form internal call records in standard format. Business analysis module: The enumerated values of each field in the call record are combined through delimited fields to form an expression. Each call record will have a unique expression corresponding to it. Through this expression, various business call record scenarios are abstracted into one expr_id value after being processed by business analysis, and then filled into the detailed list condition_id field for output, and then passed to the subsequent rating process for invocation. The duplicate check module stores the full amount of call record index information through the duplicate check MDB (memory database), and provides an interface for prior call record index query and update. The rating module combines the user's subscription information and rates the call records according to the agreed rates to complete the update of the charging and settlement of the user's original call records.
[0076] As the types and volumes of operator services increase, the traditional call record duplicate check method cannot meet the call record rating mode in the existing high-concurrency scenarios, and it is easy to have duplicate check backlogs caused by the inability of concurrency to take effect. That is, in the related technology, when duplicate check on call records is required, usually artificial selection of duplicate check indexes is made from a large number of call record fields, and the optimal key field combination cannot be automatically selected according to the actual call record situation, resulting in low call record duplicate check efficiency. Moreover, the duplicate check fields designed manually cannot reach the optimal state, with redundancy or lack. Among them: redundant duplicate check fields will occupy the memory database and also lead to low duplicate check efficiency. The lack of duplicate check fields will result in insufficient duplicate check accuracy, ultimately leading to duplicate rating and causing customer complaints about charging or settlement.
[0077] To solve the above defects existing in the related technology, this article proposes a call record duplicate check method. From the perspective of improving the call record duplicate check efficiency and rating accuracy, it can optimize the selection of key index fields on the premise of a certain memory, and achieve the accuracy of duplicate check and the timeliness of call record processing.
[0078] Hereinafter, the call record duplicate check method proposed by the present invention will be further explained through specific embodiments.
[0079] In one embodiment, please refer to Figure 2 , the call record duplicate check method includes the following steps:
[0080] Step S10: Obtain the keyword fields and preset field combinations associated with the call records to be checked for duplicates in the current period;
[0081] Step S20: Generate an initial keyword field combination based on the preset field combination and the keyword fields;
[0082] Step S30: Determine the risk coefficient corresponding to each initial keyword field combination according to a preset keyword field duplicate-checking risk coefficient function model;
[0083] Step S40: Select a target keyword field combination from the initial keyword field combinations according to the risk coefficient;
[0084] Step S50: Perform duplicate-checking on the call records to be checked for duplicates in the current period based on the target keyword field combination.
[0085] In this embodiment, the time limit period corresponding to duplicate-checking analysis can be set. For example, when the business volume is large, the expiration period can be set to 1 day. When the business volume is small, the time limit period can also be set to 1 week. Or, according to the cache and / or computing power of the system, the time limit period can be set to N days. The period length of the time limit period can be custom-set according to system requirements. This embodiment does not make specific limitations in this regard.
[0086] Furthermore, the keyword fields and preset field combinations associated with the call records to be checked for duplicates in the current period can be obtained. It can be understood that the keyword fields associated with the call records to be checked for duplicates can refer to all keyword fields associated with the call records to be checked for duplicates. The preset field combination is preset and includes a field combination of one or more keyword fields. Among them, the preset field sub is a custom field combination. For example, for a roaming traffic CNGO file, the following can be selected: user number | start time | roaming traffic value as its corresponding preset field combination.
[0087] After obtaining the keyword fields and preset field combinations associated with the call records to be checked for duplicates in the current period, based on the preset field combination, multiple initial keyword field combinations can be generated according to the keyword fields associated with the call records to be checked for duplicates. That is, on the basis of the preset combined fields, each time one or more new keyword fields are selected from the keyword fields associated with the call records to be checked for duplicates and added to the preset combined fields to form an initial keyword field combination.
[0088] Exemplarily, taking the preset combined field X 0 as the basis, in the keyword fields associated with the call records to be checked for duplicates, a new keyword field that is not included in X 0 is selected and combined with X 0 to form an initial keyword field combination X 1 , and then taking X1 Based on this, select an X 1 New keyword fields not included in and X 1 Form the initial keyword field combination X 2 And so on until the initial keyword field combination X containing all keyword fields is generated n .
[0089] Furthermore, after determining the initial keyword field combination, the risk coefficient corresponding to each initial keyword field combination can be determined according to a preset keyword field duplicate check risk coefficient function model.
[0090] It can be understood that, as an implementation, steps S20 and S30 can also be executed simultaneously. That is, when generating each initial keyword field combination, the risk coefficient corresponding to the initial keyword field combination can be determined.
[0091] Furthermore, after determining the risk coefficients corresponding to each initial keyword field combination, the optimal keyword field combination can be selected from the initial keyword field combinations as the target keyword field combination according to the risk coefficient.
[0092] Exemplarily, the optimal keyword fields for the current cycle can be confirmed based on minimizing the risk coefficient. For example, the mean σ and standard deviation μ of the risk coefficients corresponding to each initial keyword field combination can be taken one by one, and the coefficient of variation C.V is calculated according to the following formula
[0093] C.V = σ / μ
[0094] Furthermore, based on the coefficient of variation, the initial keyword field combination with the minimum risk coefficient is determined as the target keyword field combination. Furthermore, based on the target keyword field combination, duplicate check of the call records to be checked in the current cycle is performed.
[0095] Optionally, when performing duplicate check of the call records to be checked in the current cycle based on the target keyword field combination, the first call record information digest corresponding to the prior call record and the second call record information digest corresponding to the call record to be checked can be determined first based on the target keyword field combination, and then it is determined whether the call record to be checked is a duplicate call record according to the first call record information digest and the second call record information digest.
[0096] It should be noted that when determining the first call record information digest corresponding to the prior call record and the second call record information digest corresponding to the call record to be checked for duplication based on the target keyword field combination, the index information corresponding to the prior call record can be determined first according to the target keyword field combination, and then the index information is used as the input parameter of the Secure Hash Algorithm 1 (SHA1). The string corresponding to the index information is calculated through the SHA1 algorithm, and the string is used as the first call record information digest; and based on the target keyword field combination, the second call record information digest corresponding to the call record to be checked for duplication is calculated through the SHA1 algorithm.
[0097] Exemplarily, the index information corresponding to the target keyword field combination of the prior call record can be extracted as the input information, and then the input information is processed. Among them, processing the input information may include filling the length information of the information message, information packet processing, and initializing the cache. Since SHA1 uses a 160-bit information digest, 5 link variables are required.
[0098] Then, based on the SHA1 algorithm, the string corresponding to the index information is calculated, that is, the first call record information digest. And the first call record information digest of the prior call record in the current cycle is stored through the duplicate check MDB in-memory database.
[0099] Furthermore, for the newly input call record, based on the same method above, the second call record information digest corresponding to it is calculated through the SHA1 algorithm. According to the user profile MDB, the user call records that do not need to be priced are excluded, and the call record information digests of the remaining newly input call records and the prior call records in the MDB are compared (that is, comparing the first call record information digest with the second call record information digest). Then, the duplicate call records and non-duplicate call records are determined according to the comparison result.
[0100] Optionally, after determining the duplicate call records and non-duplicate call records, the non-duplicate call records can be used as normal call records, and the normal call records are priced. Then, the pricing amount of the priced normal call records is obtained, and based on the pricing amount and the target keyword field combination, the priced normal call records are checked for duplicate call records to determine whether there are duplicate priced call records. That is, the pricing amount and the target keyword field combination are combined into a new review keyword field combination, and then the priced call records are checked for duplication based on the review keyword field combination. Thus, it is further determined whether there are duplicate priced call records.
[0101] Optionally, in order to prevent the concurrent processing from failing when processing a high volume of service call records, a token bucket mechanism is used for processing the priced call records. The token bucket method refers to filling the internal storage pool at a given rate, and the token is a virtual information packet. The work process includes:
[0102] a. Put the call detail record files to be processed into the token bucket - internal storage pool at a specific rate, classify the call detail records according to the matching rules, and set the call detail record priority at the same time;
[0103] b. The call detail record files that meet the matching rules enter the token bucket storage pool for processing. When there are enough tokens in the storage pool, they can be continuously processed and stored in the database, and the amount of tokens in the token bucket decreases correspondingly according to the size of the file;
[0104] c. When there are insufficient tokens in the storage pool, they will be processed according to the file priority sorting. At the same time, when the system concurrent processing fails, in order to prevent file backlog, the system allows burst sending to achieve the purpose of stable processing.
[0105] Optionally, when it is determined that the call detail record is a duplicate call detail record according to the comparison result of the first call detail record information digest and the second call detail record information digest, it is also possible to audit and back - check the call detail records with the same information digest in the cache MDB to add keyword fields for comparison. If there is an inconsistency after adding the fields, a reverse recovery rating request is initiated to perform rating processing on the misjudged call detail record.
[0106] Optionally, referring to Figure 4 , as an optional implementation solution, before step S30, it further includes:
[0107] Step S60: Obtain the first parameter estimation result and the first system observation information corresponding to the previous cycle;
[0108] Step S70: Calculate the Kalman gain according to the first parameter estimation result and the first system observation information;
[0109] Step S80: Determine the optimization parameter according to the Kalman gain and the system measurement parameter corresponding to the current cycle;
[0110] Step S90: Optimize the parameters of the preset keyword field duplicate check risk coefficient function model according to the optimization parameter
[0111] In this implementation solution, the first parameter estimation result and the first system observation information corresponding to the previous cycle can be obtained first, and then the Kalman gain K is calculated according to the following formula T :
[0112]
[0113] Among them, is the parameter estimation result of the previous cycle θ T-1 (that is, the covariance matrix of the first parameter estimation result), is the system observation information of the previous cycle (that is, the first - order derivative of the first system observation information); the covariance of the measurement noise.
[0114] Further, the current Kalman gain K is determined T After that, the system measurement parameter corresponding to the current period can be calculated according to the following formula to determine the optimization parameter θ T :
[0115]
[0116] wherein is the system measurement parameter corresponding to the current period, is the first-order derivative of θT.
[0117] It can be understood that, in order to construct the Kalman filter and realize the autoregressive update of the prior demand parameters to the risk parameters of the next stage step by step for the purpose of autoregressive update, the covariance matrix P of the parameter estimation result corresponding to the next period can also be calculated according to the following formula T and the system observation information Z T .
[0118]
[0119]
[0120] where is the measurement noise, is the first-order derivative of PT.
[0121] In the technical solution disclosed in this embodiment, first, the keyword fields and the preset field combinations associated with the call records to be checked for duplicates in the current period are obtained, then the initial keyword field combinations are generated based on the preset field combinations and the keyword fields, and then the risk coefficient corresponding to each initial keyword field combination is determined according to the preset keyword field duplicate-checking risk coefficient function model. The target keyword field combination is selected from the initial keyword field combinations according to the risk coefficient, and finally, the call records to be checked for duplicates in the current period are checked for duplicates based on the target keyword field combination. Since the optimal keyword field combination can be automatically selected according to the actual call record situation, the effect of improving the efficiency of call record duplicate-checking can be achieved.
[0122] In addition, an embodiment of the present invention also provides a call record duplicate-checking device, which includes: a memory, a processor, and a call record duplicate-checking program stored on the memory and executable on the processor. When the call record duplicate-checking program is executed by the processor, the steps of the call record duplicate-checking method described in each of the above embodiments are implemented.
[0123] In addition, please refer to Figure 5 , an embodiment of the present invention also provides a call record duplicate-checking device 100, which includes:
[0124] An acquisition module 101, configured to acquire keyword fields and a preset field combination associated with the call records to be checked for duplication in the current period;
[0125] A generation module 102, configured to generate an initial keyword field combination based on the preset field combination and the keyword fields;
[0126] A calculation module 103, configured to determine a risk coefficient corresponding to each of the initial keyword field combinations according to a preset keyword field duplication risk coefficient function model;
[0127] A selection module 104, configured to select a target keyword field combination from the initial keyword field combinations according to the risk coefficient;
[0128] A duplication check module 105, configured to perform call record duplication check on the call records to be checked for duplication in the current period based on the target keyword field combination.
[0129] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a call record duplication check program is stored. When the call record duplication check program is executed by a processor, the steps of the call record duplication check method described in each of the above embodiments are implemented.
[0130] It should be noted that, in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.
[0131] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0132] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) as described above, and includes several instructions for causing a call record duplication check device (such as a server) to execute the methods described in each embodiment of the present invention.
[0133] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A call record duplicate checking method, characterized in that, the call record duplicate checking method includes: Obtaining the key fields associated with the call records to be checked for duplicates in the current period and a preset field combination, where the key fields associated with the call records to be checked for duplicates are all the key fields associated with the call records to be checked for duplicates, and the preset field combination includes a field combination of one or more key fields; Generating a plurality of initial key field combinations based on the preset field combination and the key fields. On the basis of the preset combination field, each time one or more new key fields are selected from the key fields associated with the call records to be checked for duplicates and added to the preset combination field to form an initial key field combination; Determining the risk coefficient corresponding to each of the initial key field combinations according to a preset key field duplicate checking risk coefficient function model; Determining the initial key field combination with the smallest risk coefficient as the target key field combination; Performing call record duplicate checking on the call records to be checked for duplicates in the current period based on the target key field combination.
2. The call record duplicate checking method according to claim 1, characterized in that, before the step of determining the risk coefficient corresponding to each of the initial key field combinations according to a preset key field duplicate checking risk coefficient function model, it further includes: Obtaining the first parameter estimation result and the first system observation information corresponding to the previous period; Calculating the Kalman gain according to the first parameter estimation result and the first system observation information; Determining the optimization parameter according to the Kalman gain and the system measurement parameters corresponding to the current period; Performing parameter optimization on the preset key field duplicate checking risk coefficient function model according to the optimization parameter.
3. The call record duplicate checking method according to claim 2, characterized in that, after the step of calculating the Kalman gain according to the first parameter estimation result and the first system observation information, it further includes: Determining the second parameter estimation result and the second system observation information corresponding to the next period according to the Kalman gain and the system measurement parameters corresponding to the current period, and saving the second parameter estimation result and the second system observation information.
4. The call record duplicate checking method according to claim 1, characterized in that, the step of performing call record duplicate checking on the call records to be checked for duplicates in the current period based on the target key field combination includes: Determining the first call record information digest corresponding to the prior call record and the second call record information digest corresponding to the call record to be checked for duplicates based on the target key field combination; Determining whether the call record to be checked for duplicates is a duplicate call record according to the first call record information digest and the second call record information digest.
5. The call record duplicate checking method according to claim 4, characterized in that, the step of determining the first call record information digest corresponding to the prior call record and the second call record information digest corresponding to the call record to be checked for duplicates based on the target key field combination includes: Determining the index information corresponding to the prior call record according to the target key field combination; Use the cable information as an input parameter of the Secure Hash Algorithm SHA1, calculate the string corresponding to the cable information through the SHA1 algorithm, and use the string as the first call record information digest; and Based on the target keyword field combination, calculate the second call record information digest corresponding to the call record to be checked for duplication through the SHA1 algorithm.
6. The call record duplication checking method according to claim 1,[[]] characterized in that,[[]] after the step of checking for duplication of the call record to be checked for duplication corresponding to the current period based on the target keyword field combination, further includes:[[]] Determine normal call records according to the call record duplication checking result, and perform rating processing on the normal call records; Obtain the rating amount of the normal call records after the rating processing; Perform call record duplication checking on the normal call records after the rating processing according to the rating amount and the target keyword field combination to determine whether there are duplicate rated call records.
7. The call record duplication checking method according to claim 6,[[]] characterized in that,[[]] When performing rating processing on the normal call records, a token bucket mechanism is used for processing rated call records.
8. A call record duplication checking device,[[]] characterized in that,[[]] The call record duplication checking device includes: a memory, a processor, and a call record duplication checking program stored on the memory and executable on the processor. When the call record duplication checking program is executed by the processor, the steps of the call record duplication checking method according to any one of claims 1 to 7 are implemented.
9. A call record duplication checking device,[[]] characterized in that,[[]] The call record duplication checking device includes:[[]] An acquisition module, configured to acquire keyword fields associated with the call record to be checked for duplication corresponding to the current period and a preset field combination. The keyword fields associated with the call record to be checked for duplication are all keyword fields associated with the call record to be checked for duplication, and the preset field combination includes a field combination of one or more keyword fields; A generation module, configured to generate a plurality of initial keyword field combinations based on the preset field combination and the keyword fields. On the basis of the preset combination field, each time one or more new keyword fields are selected from the keyword fields associated with the call record to be checked for duplication and added to the preset combination field to form an initial keyword field combination; A calculation module, configured to determine the risk coefficient corresponding to each initial keyword field combination according to a preset keyword field duplication checking risk coefficient function model; A selection module, configured to determine the initial keyword field combination with the smallest risk coefficient as the target keyword field combination; A duplication checking module, configured to perform call record duplication checking on the call record to be checked for duplication corresponding to the current period based on the target keyword field combination.
10. A computer-readable storage medium,[[]] characterized in that,[[]] A call record duplication checking program is stored on the computer-readable storage medium. When the call record duplication checking program is executed by a processor, the steps of the call record duplication checking method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and device for auditing phone bills with different sources
CN103167202A