Method for processing credit investigation data and related device

By using keyword matching values, backbone information correct values ​​and sensitive values ​​in the credit data processing method, and performing data cleaning, the problem of unstable credit data quality is solved and the data quality and evaluation accuracy is improved.

CN115168360BActive Publication Date: 2025-06-20SHENZHEN WEIZHONG TAXATION INFORMATION SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210846619.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-06-20
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

In the prior art, the quality of credit reporting data is uneven, making it difficult to meet the users' needs for risk assessment based on credit reporting data.

Method used

By obtaining the matching values ​​of multiple keywords and main information in the credit report data, the correct values ​​and sensitive values ​​of the main information in the credit report data, the quality value of the credit report data is determined, and cleaning it according to the preset quality value, to improve the quality of the credit report data.

Benefits of technology

It improves the accuracy of judging the quality of credit reporting data, improves the quality of credit reporting data obtained, and meets users' needs for high-quality credit reporting data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115168360B_ABST
    Figure CN115168360B_ABST
Patent Text Reader

Abstract

The present application provides a method for processing credit investigation data and related devices, including obtaining a first matching value between each keyword in multiple keywords in the credit investigation data and the backbone information, obtaining a second matching value for characterizing the correlation degree between the backbone information and the credit investigation data according to the first matching value of the keyword, obtaining a backbone information correct value for characterizing the quantity of the correct backbone information, obtaining a sensitivity value for characterizing the quantity of sensitive words in the credit investigation data according to the quantity of sensitive words in the credit investigation data, determining a quality value for characterizing the quality level of the credit investigation data according to the second matching value, the backbone information correct value, and the sensitivity value, and if the quality value does not reach a preset quality value, cleaning the credit investigation data according to a preset data cleaning method. In the present application, the quality of the credit investigation data is judged by determining whether the obtained quality value reaches the preset quality value, so as to improve the accuracy of judging the quality of the credit investigation data and obtain credit investigation data meeting the required quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data processing, and particularly relates to a method for processing credit investigation data and related devices. Background Art

[0002] Credit investigation data, also known as credit information, is called enterprise credit information when reflecting the credit status of an enterprise, and is called personal credit information when reflecting the credit status of an individual.

[0003] If there are abnormal parts in the credit investigation data of an enterprise, professional personnel of financial institutions need to investigate the enterprise in a timely manner to check whether the enterprise has the ability to repay the loan. In the prior art, due to the numerous sources of credit investigation data, the quality of the obtained credit investigation data is uneven, making it difficult to meet the user's need for risk assessment based on credit investigation data. Summary of the Invention

[0004] This application provides a method for processing credit investigation data and related devices. By determining whether the quality value of the obtained credit investigation data reaches a preset quality value, the quality of the credit investigation data is determined, the accuracy of judging the quality of the credit investigation data is improved, and the quality of the obtained credit investigation data is improved.

[0005] In a first aspect, this application provides a method for processing credit investigation data, including:

[0006] Obtain a first matching value between each keyword in the credit investigation data and the backbone information, where the first matching value is used to represent the degree of association between the backbone information and the keyword, and the backbone information is used to represent the key information in the credit investigation data;

[0007] Obtain a second matching value according to the first matching value of the keyword in the multiple keywords, where the second matching value is used to represent the degree of association between the backbone information and the credit investigation data;

[0008] Obtain the correct value of the backbone information, where the correct value of the backbone information is used to represent the amount of correct backbone information in the backbone information;

[0009] Obtain a sensitivity value according to the number of sensitive words in the credit investigation data, where the sensitivity value is used to represent the number of sensitive words in the credit investigation data;

[0010] Determine the quality value of the credit investigation data according to the second matching value, the correct value of the backbone information, and the sensitivity value, where the quality value is used to represent the quality level of the credit investigation data;

[0011] Judge whether the quality value reaches a preset quality value;

[0012] Otherwise, clean the credit investigation data according to a preset data cleaning method so that the quality value of the credit investigation data reaches a preset quality value.

[0013] In a second aspect, the present application provides a processing device for credit investigation data, including:

[0014] A first acquisition unit, configured to acquire a first matching value between each keyword in multiple keywords in the credit investigation data and the backbone information, where the first matching value is used to represent the correlation degree between the backbone information and the keyword, and the backbone information is used to represent key information in the credit investigation data;

[0015] A second acquisition unit, configured to acquire a second matching value according to the first matching value of the keyword in the multiple keywords, where the second matching value is used to represent the correlation degree between the backbone information and the credit investigation data;

[0016] A third acquisition unit, configured to acquire the correct value of the backbone information, where the correct value of the backbone information is used to represent the amount of correct backbone information in the backbone information;

[0017] A fourth acquisition unit, configured to acquire a sensitivity value according to the number of sensitive words in the credit investigation data, where the sensitivity value is used to represent the number of sensitive words in the credit investigation data;

[0018] A determination unit, configured to determine the quality value of the credit investigation data according to the second matching value, the correct value of the backbone information, and the sensitivity value, where the quality value is used to represent the quality level of the credit investigation data;

[0019] A judgment unit, configured to judge whether the quality value reaches a preset quality value;

[0020] A processing unit, configured to clean the credit investigation data according to a preset data cleaning method so that the quality value of the credit investigation data reaches a preset quality value.

[0021] In a third aspect, the present application provides an electronic device, including:

[0022] One or more processors;

[0023] One or more memories for storing programs,

[0024] The one or more memories and the program are configured to control the electronic device by the one or more processors to execute instructions of any method in the first aspect or the steps in the second aspect of the embodiments of the present application.

[0025] Fourthly, the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program for electronic data exchange. The computer program enables a computer to execute some or all of the steps described in any method of the first aspect or the second aspect of the embodiments of the present application.

[0026] Fifthly, the present application provides a computer program. The computer program is operable to enable a computer to execute some or all of the steps described in any method of the first aspect or the second aspect of the embodiments of the present application. The computer program can be a software installation package.

[0027] It can be seen that in the embodiments of the present application, the quality value of the credit investigation data is determined by obtaining the second matching value, the correct value of the backbone information, and the sensitive value of the credit investigation data. The quality value is used to characterize the quality level of the credit investigation data. The second matching value is used to characterize the correlation degree between the backbone information and the credit investigation data. The second matching value is obtained according to the first matching value between each keyword in multiple credit investigation data and the backbone information. The correct value of the backbone information is used to characterize the amount of correct backbone information in the backbone information. The sensitive value is used to characterize the number of sensitive words in the credit investigation data. By determining whether the quality value reaches a preset quality value, it is determined whether the credit investigation data is qualified. If the preset quality value is not reached, the credit investigation data is cleaned according to a preset data cleaning method so that the quality value of the credit investigation data reaches the preset quality value. Determining the quality value of the credit investigation data through the second matching value, the correct value of the backbone information, and the sensitive value improves the accuracy of the determined quality value and further improves the quality of the obtained credit investigation data. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 is a schematic structural diagram of a server provided by an embodiment of the present application;

[0030] Figure 2 is a schematic diagram of a processing system for credit investigation data provided by an embodiment of the present application;

[0031] Figure 3 is a flowchart of a method for processing credit investigation data provided by an embodiment of the present application;

[0032] Figure 4 is a block diagram of the functional units of a processing device for credit investigation data provided by an embodiment of the present application. Detailed implementation manners

[0033] In order to enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this application.

[0034] The terms "first", "second", etc. in the description and claims of this application and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0035] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0036] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a server provided by an embodiment of this application. As Figure 1 shown, the server includes one or more processors 120, a memory 130, a communication module 140, and one or more programs 131. The processor 120 is communicatively connected to the memory 130 and the communication module 140 through an internal communication bus.

[0037] Among them, the one or more programs 131 are stored in the above-mentioned memory 130 and are configured to be executed by the above-mentioned processor 120. The one or more programs 131 include instructions for executing any step in the following method embodiments.

[0038] Among them, the processor 120 can be, for example, a Central Processing Unit (CPU), a general-purpose processor, a Digital Signal Processor (DSP), an Application-Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, units, and circuits described in connection with the disclosure of this application. The processor 120 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The communication unit can be the communication module 140, a transceiver, a transceiver circuit, etc., and the storage unit can be the memory 130.

[0039] The memory 130 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus random access memory (DR RAM).

[0040] To better understand the technical solutions of the embodiments of this application, the processing system for credit investigation data that may be involved in the embodiments of this application will be introduced first.

[0041] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a credit investigation data processing system provided by an embodiment of the present application. The credit investigation data processing system includes a server and multiple terminal devices. Among them, each of the multiple terminal devices is network-connected to the server so that each terminal device can perform data interaction with the server through the network. As Figure 2 shown, it may specifically include terminal device 210a, terminal device 210b, terminal device 210c,..., terminal device 210n.

[0042] Among them, each of the multiple terminal devices is a device corresponding to a desktop that can provide voice and / or data information to participating users. The terminal device can be an intelligent terminal such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a wearable device, a head-mounted device, a vehicle-mounted terminal, etc. The specific types of terminal devices are not limited here. It should be understood that each terminal device in the credit investigation data processing system as Figure 2 shown can be installed with a credit investigation data processing application (i.e., an application client). When the application client runs on each terminal device, it can perform data interaction with the above-mentioned Figure 2 shown server respectively.

[0043] Please refer to Figure 3 , Figure 3 which is a flowchart of a credit investigation data processing method provided by an embodiment of the present application. Next, the credit investigation data processing method applied to the server involved in the embodiments of the present application will be described in detail with reference to Figure 3 . As Figure 3 shown, a credit investigation data processing method includes:

[0044] Step 310: Obtain a first matching value between each keyword in the credit investigation data and the backbone information. The first matching value is used to characterize the association degree between the backbone information and the keyword, and the backbone information is used to characterize the key information in the credit investigation data.

[0045] Specifically, after the server obtains at least one credit investigation data, it extracts the keywords of each credit investigation data in the at least one credit investigation data respectively, and extracts the main information of the credit investigation data, and calculates the first matching value between the keyword corresponding to each credit investigation data and the main information. For example, if a credit investigation data obtained includes: Xiaoming's ID number is, and the phone number is 1234567. Extract the first keyword in the credit investigation data: Xiaoming; the second keyword: ID number; the third keyword: phone number. Extract the main information: Xiaoming's phone number is 1234567. Then the first matching value between the first keyword and the main information is 13, the first matching value between the second keyword and the main information is 0, and the first matching value between the third keyword and the main information is 27. The size of the first matching value can be determined according to the proportion of the keyword in the main information, or it can also be determined according to the proportion of the synonym of the keyword in the main information. It can be understood that the determination method of the first matching value can be set according to the actual situation, and specific limitations are not made here.

[0046] Step 320: Obtain a second matching value according to the first matching value of the keyword among the multiple keywords. The second matching value is used to characterize the correlation degree between the main information and the credit investigation data.

[0047] Specifically, after obtaining the first matching value corresponding to each keyword among the multiple keywords, obtain the second matching value according to the first matching values corresponding to all the keywords to characterize the correlation degree between the main information and the credit investigation data.

[0048] Step 330: Obtain the correct value of the main information. The correct value of the main information is used to characterize the amount of the correct main information in the main information.

[0049] Specifically, obtain the correct value of the main information that is used to characterize the amount of the correct main information in the main information.

[0050] Step 340: Obtain a sensitivity value according to the number of sensitive words in the credit investigation data. The sensitivity value is used to characterize the number of sensitive words in the credit investigation data.

[0051] Specifically, count the number of sensitive words in the credit investigation data, and obtain the sensitivity value according to the number of sensitive words in the credit investigation data. The sensitive words can be pre-set words. After obtaining the credit investigation data, traverse the credit investigation data to determine whether the pre-set sensitive words are included in the credit investigation data.

[0052] It can be understood that the execution order of steps 320 to 340 can be adjusted according to the actual situation, and the execution order of steps 320 to 340 is not limited here.

[0053] Step 350: Determine the quality value of the credit investigation data according to the second matching value, the correct value of the backbone information, and the sensitive value. The quality value is used to represent the quality level of the credit investigation data.

[0054] Specifically, determine the quality value of the credit investigation data according to the obtained second matching value, the correct value of the backbone information, and the sensitive value to determine the quality level of the credit investigation data.

[0055] Step 360: Determine whether the quality value reaches a preset quality value.

[0056] Step 370: Clean the credit investigation data according to a preset data cleaning method, so that the quality value of the credit investigation data reaches the preset quality value.

[0057] Specifically, determine whether the obtained quality value reaches the preset quality value. If it reaches the preset quality value, the credit investigation data meets the quality requirements, output a quality report, and save the credit investigation data or directly use it for subsequent data processing. If it does not reach the preset quality value, clean the credit investigation data according to the preset data cleaning method so that the quality value of the credit investigation data reaches the preset quality value.

[0058] It can be seen that in this example, the quality of the credit investigation data is judged by the second matching value, the correct value of the backbone information, and the sensitive value of the credit investigation data, improving the accuracy of the judgment result. Moreover, the credit investigation data with a quality value not reaching the preset quality value is cleaned to obtain credit investigation data reaching the preset quality value, thereby improving the quality of the obtained credit investigation data and the reliability of the subsequent processing result based on the credit investigation data.

[0059] In a possible example, the obtaining of the second matching value according to the first matching value of the keyword among the multiple keywords includes: determining a preset weight value corresponding to each keyword according to the magnitude of the first matching value of the keyword among the multiple keywords; calculating the product of each first matching value and the preset weight value corresponding to the first matching value; adding the products corresponding to each first matching value to obtain the second matching value.

[0060] In a specific example, the first matching value of each keyword among multiple keywords is obtained, and the corresponding preset weight value is determined according to the magnitude of the first matching value of each keyword. After obtaining the preset weight value of each first matching value, the product of the first matching value and the preset weight value corresponding to the first matching value is calculated, and the products corresponding to each first matching value are added together to obtain the second matching value. For example, the credit investigation data includes a first keyword, a second keyword, and a third keyword. The first matching value of the first keyword is 13, the first matching value of the second keyword is 0, and the first matching value of the third keyword is 27. Calculate the total value of the first keyword, the second keyword, and the third keyword, and determine the preset weight value according to the ratio of the first matching value of each of the first keyword, the second keyword, and the third keyword to the total value. The larger the first matching value, the larger the corresponding preset weight value. Therefore, it is determined that the preset weight value corresponding to the first keyword is 32%, the preset weight value corresponding to the second keyword is 0, and the preset weight value corresponding to the third keyword is 68%. The second matching value is 13 * 32% + 27 * 68% = 22.52.

[0061] It can be seen that in this example, the product of the first matching value corresponding to each keyword and the preset weight value is obtained, and then the products are added together to obtain the second matching value, intuitively expressing the correlation degree between the main information and the credit investigation data through numerical values.

[0062] In a possible example, the first matching values corresponding to each keyword can be directly added together, and the sum is divided by the total number of keywords among the multiple keywords to obtain the average value, and this average value is the second matching value. By directly adding and then taking the average, the calculation is simplified and time is saved.

[0063] In a possible example, obtaining the correct value of the main information includes: obtaining the similarity between the main information and the preset main information; determining the correct value of the main information of the credit investigation data according to the similarity.

[0064] In a specific example, different preset main information is preset according to different quality values required by the credit investigation data. According to the quality value required by the credit investigation data selected by the user, the preset main information is determined, the similarity between the main information and the preset main information is obtained, and the correct value of the main information of the corresponding credit investigation data is determined according to the similarity. The determination process of this similarity can be: comparing the number of the same words or synonyms in the main information and the preset main information, and determining the similarity according to the number of the same or similar words.

[0065] It can be seen that in this example, by comparing the similarity between the main information of the credit investigation data and the preset main information to determine the correct value of the main information, the obtained correct value of the main information is more accurate, and the credibility of the quality value of the obtained credit investigation data is improved.

[0066] In a possible example, obtaining a sensitivity value based on the number of sensitive words in the credit investigation data includes: counting the number of sensitive words in the credit investigation data; obtaining a plurality of preset quantity ranges, where each quantity range in the plurality of preset quantity ranges corresponds to a sensitivity value; determining the preset quantity range to which the number of sensitive words belongs, and obtaining the sensitivity value.

[0067] In a specific example, count the number of sensitive words in the credit investigation data. The sensitive words can be pre-set words. After obtaining the credit investigation data, traverse the credit investigation data to determine whether it contains pre-set sensitive words. If it contains, count the number of sensitive words. Obtain a plurality of preset quantity ranges, where each numerical range corresponds to a sensitivity value, and judge the numerical range to which the obtained number of sensitive words belongs, and then determine the sensitivity value of the credit investigation data.

[0068] It can be seen that in this example, the sensitivity value is quickly determined through the number of sensitive words in the credit investigation data, saving time.

[0069] In a possible example, determining the quality value of the credit investigation data according to the second matching value, the correct value of the backbone information, and the sensitivity value includes: obtaining a preset first weight value corresponding to the second matching value; obtaining a preset second weight value corresponding to the correct value of the backbone information; obtaining a preset third weight value corresponding to the sensitivity value; determining the quality value of the credit investigation data according to a first numerical value calculated from the first weight value and the second matching value, a second numerical value calculated from the second weight value and the correct value of the backbone information, and a third numerical value obtained from the third weight value and the sensitivity value.

[0070] In a specific example, after obtaining the second matching value, the correct value of the backbone information, and the sensitivity value, respectively obtain a preset first weight value corresponding to the second matching value, a preset second weight value corresponding to the correct value of the backbone information, and a preset third weight value corresponding to the sensitivity value. It can be understood that the weight values corresponding to the second matching value, the correct value of the backbone information, and the sensitivity value can be preset according to the actual needs of the user, or the corresponding weight values can be determined respectively according to the magnitudes of the second matching value, the correct value of the backbone information, and the sensitivity value. There is no specific limitation here. Determine the quality value of the credit investigation data according to a first numerical value calculated from the first weight value and the second matching value, a second numerical value calculated from the second weight value and the correct value of the backbone information, and a third numerical value calculated from the third weight value and the sensitivity value.

[0071] It can be seen that in this example, the quality value of the credit investigation data is determined according to the second matching value, the correct value of the backbone information, and the sensitivity value, improving the accuracy of the determination result.

[0072] In a possible example, after cleaning the credit investigation data according to the preset data cleaning method, the following steps are further included: sending a notification message to the user's terminal device, where the notification message is used to remind the user that the credit investigation data that does not reach the preset quality value has been cleaned up.

[0073] In a specific example, after cleaning the credit investigation data according to the preset data cleaning method, a notification message is sent to the user's terminal device. The notification message includes information on the credit investigation data that does not reach the preset quality value, and the notification message is used to remind the user that the credit investigation data that does not reach the preset quality value has been cleaned up.

[0074] It can be seen that in this example, a notification message is sent to the user's terminal device to remind the user that the credit investigation data has been cleaned up, avoiding real-time viewing during the cleaning process, saving time, and improving the user experience.

[0075] In a possible example, the preset data cleaning method includes, but is not limited to, deletion, duplicate reduction, and desensitization processing.

[0076] Specifically, the preset data cleaning method includes, but is not limited to, deletion, duplicate reduction, and desensitization processing. For example, the preset data cleaning method may further include supplementing the credit investigation data.

[0077] The embodiments of the present application may divide the functional units of the processing device for credit investigation data according to the above method examples. For example, each functional unit may be divided corresponding to each function, or two or more functions may be integrated into one processing unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit. The division of units in the embodiments of the present application is illustrative, and is only a logical functional division. There may be other division methods in actual implementation.

[0078] In the case of dividing each functional unit corresponding to each function, please refer to Figure 4 , Figure 4 which is a block diagram of the functional unit composition of a processing device for credit investigation data provided by the embodiments of the present application, and includes:

[0079] A first acquisition unit 410, configured to acquire a first matching value between each keyword in the credit investigation data and the backbone information, where the first matching value is used to represent the degree of association between the backbone information and the keyword, and the backbone information is used to represent the key information in the credit investigation data;

[0080] A second acquisition unit 420, configured to acquire a second matching value according to the first matching value of the keyword among the multiple keywords, where the second matching value is used to represent the degree of association between the backbone information and the credit investigation data;

[0081] A third acquisition unit 430, configured to acquire the correct value of the backbone information, where the correct value of the backbone information is used to characterize the amount of correct backbone information in the backbone information;

[0082] A fourth acquisition unit 440, configured to acquire a sensitivity value according to the number of sensitive words in the credit investigation data, where the sensitivity value is used to characterize the number of sensitive words in the credit investigation data;

[0083] A determination unit 450, configured to determine a quality value of the credit investigation data according to the second matching value, the correct value of the backbone information, and the sensitivity value, where the quality value is used to characterize the quality level of the credit investigation data;

[0084] A judgment unit 460, configured to judge whether the quality value reaches a preset quality value;

[0085] A processing unit 470, configured to clean the credit investigation data according to a preset data cleaning method, so that the quality value of the credit investigation data reaches the preset quality value.

[0086] In a possible example, the second acquisition unit 420 is further configured to: determine a preset weight value corresponding to each keyword according to the magnitude of the first matching value of the keyword among the multiple keywords; calculate the product of each first matching value and the preset weight value corresponding to the first matching value; and add up the products corresponding to each first matching value to obtain the second matching value.

[0087] In a possible example, the third acquisition unit 430 is further configured to acquire the similarity between the backbone information and a preset backbone information; and determine the correct value of the backbone information of the credit investigation data according to the similarity.

[0088] In a possible example, the fourth acquisition unit 440 is further configured to count the number of sensitive words in the credit investigation data; acquire a plurality of preset quantity ranges, where each quantity range in the plurality of preset quantity ranges corresponds to a sensitivity value; and determine the sensitivity value according to the preset quantity range to which the number of sensitive words belongs.

[0089] In a possible example, the determination unit 450 is further configured to acquire a preset first weight value corresponding to the second matching value; acquire a preset second weight value corresponding to the correct value of the backbone information; acquire a preset third weight value corresponding to the sensitivity value; and determine the quality value of the credit investigation data according to a first numerical value calculated from the first weight value and the second matching value, a second numerical value calculated from the second weight value and the correct value of the backbone information, and a third numerical value obtained from the third weight value and the sensitivity value.

[0090] In a possible example, the method further includes a sending unit configured to send a notification message to the user's terminal device, where the notification message is used to remind the user that the credit investigation data that has not reached the preset quality value has been cleaned up.

[0091] In a possible example, the preset data cleaning methods include, but are not limited to, deletion, duplicate reduction, and desensitization processing.

[0092] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0093] The embodiments of the present application also provide a computer storage medium, where the computer storage medium stores a computer program for electronic data exchange, and the computer program causes the computer to execute some or all of the steps of any of the methods described in the above method embodiments. The above computer includes an electronic device.

[0094] The embodiments of the present application also provide a computer program product, where the computer program product includes a computer program, and the computer program is operable to cause the computer to execute some or all of the steps of any of the methods described in the above method embodiments.

[0095] The computer program product can be a software installation package, and the above computer includes an electronic device.

[0096] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0097] In several embodiments provided by the present application, it should be understood that the disclosed methods, devices, and systems can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of the units is only a logical function division, and there can be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0098] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0099] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can be physically included separately, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0100] The above integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above software functional units stored in a storage medium include several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute some steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0101] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application, and these improvements and refinements are also regarded as the protection scope of the present application.

Claims

1. A method for processing credit investigation data, characterized in that, Including: Obtain the first matching value of each keyword in multiple keywords in the credit investigation data and the backbone information, where the first matching value is used to characterize the correlation degree between the backbone information and the keyword, and the backbone information is used to characterize the key information in the credit investigation data; Obtain a second matching value according to the first matching value of the keyword in the multiple keywords, where the second matching value is used to characterize the correlation degree between the backbone information and the credit investigation data; Obtain the correct value of the backbone information, where the correct value of the backbone information is used to characterize the amount of correct backbone information in the backbone information; Obtain a sensitivity value according to the number of sensitive words in the credit investigation data, where the sensitivity value is used to characterize the number of sensitive words in the credit investigation data; Determine the quality value of the credit investigation data according to the second matching value, the correct value of the backbone information, and the sensitivity value, where the quality value is used to characterize the quality level of the credit investigation data; Judge whether the quality value reaches a preset quality value; If not, clean the credit investigation data according to a preset data cleaning method so that the quality value of the credit investigation data reaches the preset quality value.

2. The method for processing credit investigation data according to claim 1, characterized in that, The obtaining the second matching value according to the first matching value of the keyword in the multiple keywords includes: Determine the preset weight value corresponding to each keyword according to the size of the first matching value of each keyword in the multiple keywords; Calculate the product of the first matching value of each keyword and the preset weight value corresponding to the keyword determined based on the size of the first matching value; Add the products of each first matching value and the corresponding preset weight value to obtain the second matching value.

3. The method for processing credit investigation data according to claim 1, characterized in that, The obtaining the correct value of the backbone information includes: Obtain the similarity between the backbone information and the preset backbone information; Determine the correct value of the backbone information of the credit investigation data according to the similarity.

4. The method for processing credit investigation data according to claim 1, characterized in that, The obtaining the sensitivity value according to the number of sensitive words in the credit investigation data includes: Count the number of sensitive words in the credit investigation data; Obtain multiple preset quantity ranges, where each quantity range in the multiple preset quantity ranges corresponds to a sensitivity value; Determine the preset quantity range to which the number of sensitive words belongs to obtain the sensitivity value.

5. The method for processing credit investigation data according to any one of claims 1-4, characterized in that, The determining the quality value of the credit investigation data according to the second matching value, the correct value of the backbone information, and the sensitivity value includes: Obtain the preset first weight value corresponding to the second matching value; Obtain the preset second weight value corresponding to the correct value of the backbone information; Obtain the preset third weight value corresponding to the sensitivity value; Determine the quality value of the credit investigation data according to the first numerical value calculated from the first weight value and the second matching value, the second numerical value calculated from the second weight value and the correct value of the backbone information, and the third numerical value obtained from the third weight value and the sensitivity value.

6. The method for processing credit investigation data according to any one of claims 1-4, characterized in that, After cleaning the credit investigation data according to the preset data cleaning method, it further includes: Send a notification message to the user's terminal device, where the notification message is used to remind the user that the credit investigation data that has not reached the preset quality value has been cleaned.

7. The method for processing credit investigation data according to any one of claims 1-4, characterized in that, The preset data cleaning method includes but is not limited to deletion, weight reduction, and desensitization processing.

8. A device for processing credit investigation data, characterized in that, Including: A first acquisition unit, configured to acquire a first matching value between each keyword in the credit investigation data and the backbone information, where the first matching value is used to characterize the association degree between the backbone information and the keyword, and the backbone information is used to characterize the key information in the credit investigation data; A second acquisition unit, configured to acquire a second matching value according to the first matching value of the keyword in the multiple keywords, where the second matching value is used to characterize the association degree between the backbone information and the credit investigation data; A third acquisition unit, configured to acquire the correct value of the backbone information, where the correct value of the backbone information is used to characterize the amount of the correct backbone information in the backbone information; A fourth acquisition unit, configured to acquire a sensitivity value according to the number of sensitive words in the credit investigation data, where the sensitivity value is used to characterize the number of sensitive words in the credit investigation data; A determination unit, configured to determine the quality value of the credit investigation data according to the second matching value, the correct value of the backbone information, and the sensitivity value, where the quality value is used to characterize the quality level of the credit investigation data; A judgment unit, configured to judge whether the quality value reaches a preset quality value; A processing unit, configured to clean the credit investigation data according to a preset data cleaning method, so that the quality value of the credit investigation data reaches the preset quality value.

9. An electronic device, characterized in that, including: A processor and a memory, where the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by the processor, the processor is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Credit risk transmission method based on credit subject association strength

    CN112529681A

  • Credit risk monitoring method for behaviors of dishonesty subjects

    CN114528460A