A method and system for enhancing object information value based on multi-source data

By extracting and merging multi-source data objects from data lakes in the government sector, and performing credibility identification and field supplementation, the problem of government data silos has been solved, the accuracy and completeness of data have been improved, and the efficient conduct of government work has been promoted.

CN115563196BActive Publication Date: 2025-12-09SHANGHAI BIG DATA INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211103709.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-09
Publication Date
2025-12-09
Estimated Expiration
2042-09-09

AI Technical Summary

Technical Problem

In the field of government affairs, multi-source data originates from different departments and systems, resulting in data silos and poor quality and correlation, which affects the accuracy and efficiency of big data analysis.

Method used

By forming a data lake, multi-source data that meets the missing field criteria are extracted as standard data objects, merged and subjected to credibility identification and missing field supplementation, and matched with related data objects to form a unique and complete credible data object.

Benefits of technology

It improved the accuracy and completeness of data, enhanced the accuracy of data application in government services, and enabled the efficient conduct of government work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115563196B_ABST
    Figure CN115563196B_ABST
Patent Text Reader

Abstract

The application provides a method and system for enhancing object information value based on multi-source data, and relates to the technical field of big data analysis, which comprises the following steps: obtaining multi-source data by each government affair system, extracting multi-source data meeting the demand field missing standard and containing business unique code as standard data objects from a plurality of preset demand fields; screening target data objects from each standard data object according to a preset business condition, and identifying the credibility of the target data objects to screen out credible data objects; supplementing missing field information of each credible data object to obtain supplemented data objects; matching multi-source data of associated business scenes according to the business unique codes of each supplemented data object to obtain matched data objects; and processing each supplemented data object and each matched data object to obtain unique complete credible data objects associated with the business unique codes. The beneficial effect is to effectively improve the accuracy and integrity of the data, so that the data has high information value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of big data analysis, and particularly relates to a method and system for enhancing object information value based on multi-source data. BACKGROUND

[0002] In recent years, big data technology and cloud computing technology have been widely applied in various fields of society. Among them, in order to improve the efficiency of government departments, the electronic government affairs system taking big data technology as the core emerges as the times require, and the electronic government affairs system integrated with big data technology has significantly improved the efficiency in data acquisition, processing, analysis and other aspects, laying a foundation for efficient government work.

[0003] However, based on the business scenarios in the government field, since the data comes from different departments and systems, the data exists in the form of an island and the data quality and correlation are poor, which directly affects the accuracy of big data analysis based on this part of data and cannot efficiently empower the business. SUMMARY

[0004] In view of the problems in the prior art, the present application provides a method for enhancing object information value based on multi-source data, comprising:

[0005] Step S1, according to the big data analysis requirements of at least one business scenario, obtaining multi-source data from each government system to form a data lake;

[0006] Step S2, for each business scenario, according to a plurality of preset requirement fields associated with the business scenario, extracting the multi-source data satisfying a requirement field missing standard and containing a corresponding business unique code from the data lake as standard data objects, and merging each standard data object with the same business unique code to form a corresponding standard object set;

[0007] Step S3, for each standard object set, according to at least one business condition associated with the business scenario, screening each standard data object to obtain a corresponding target data object, and performing credibility identification on the target data object to screen out a corresponding credible data object, to obtain a corresponding credible object set;

[0008] Step S4, for each credible object set, respectively supplementing each credible data object to obtain a corresponding supplemented data object containing all the preset requirement fields;

[0009] Step S5, respectively according to the business unique code of each supplemented data object, matching the multi-source data associated with the business scenario and excluding the standard data objects from the data lake to obtain a matching data object;

[0010] Step S6, for each of the business unique code, processing the unique complete trusted data object associated with the business unique code according to each of the supplemented data object and each of the matched data object to enhance the object information value of the multi-source data.

[0011] Preferably, the step S2 comprises:

[0012] Step S21, respectively counting the missing field quantity of each of the preset demand field missing the business scene association in each of the multi-source data;

[0013] Step S22, extracting each of the multi-source data with the least missing field quantity as the multi-source data meeting the demand field missing standard;

[0014] Step S23, extracting each of the multi-source data containing the business unique code from each of the multi-source data meeting the demand field missing standard as the standard data object, and merging each of the standard data objects with the same business unique code to form the corresponding standard object set.

[0015] Preferably, before performing the step S23, it further comprises screening each of the multi-source data meeting the demand field missing standard according to a preset time window;

[0016] Then in the step S23, the standard data object is extracted from each of the multi-source data meeting the demand field missing standard and the time window.

[0017] Preferably, the step S3 comprises:

[0018] Step S31, for each of the standard object set, respectively performing feature analysis on the field information of the preset demand field contained in each of the standard data object, and when any feature analysis result indicates that the associated business condition is met, the corresponding standard data object is reserved as the target data object;

[0019] Step S32, respectively performing standardization processing on the field information of the preset demand field contained in each of the target data object to obtain standardized field information, and performing trustworthiness identification on each of the target data object according to a corresponding preset standardization degree, so as to reserve the target data object corresponding to the standardized field information meeting the preset standardization degree as the trusted data object, and obtain the corresponding trusted object set.

[0020] Preferably, the step S4 comprises:

[0021] Step S41, for each of the trusted object set, respectively, the missing preset demand field of each of the trusted data object is counted;

[0022] Step S42, for each of the missing preset demand field, the field information of each of the trusted data object containing the preset demand field is extracted and de-duplicated respectively;

[0023] Step S43, according to the field information after de-duplication, the trusted data object with missing preset demand field is supplemented respectively, and the corresponding supplemented data object containing all the preset demand fields is obtained.

[0024] Preferably, step S6 comprises:

[0025] Step S61, for each of the supplemented preset demand field in each of the supplemented data object associated with the business unique code, the similarity between the missing field information supplemented in each of the supplemented data object and the corresponding field information of the matching data object is calculated respectively;

[0026] Step S62, the field information of the matching data object corresponding to the similarity greater than a threshold value is extracted respectively, and the extracted field information is merged and de-duplicated to obtain the merged and de-duplicated field information associated with each of the preset demand field;

[0027] Step S63, the supplemented missing field information and the merged and de-duplicated field information associated with the same business unique code and the preset demand field are added to a to-be-verified set as to-be-verified field information;

[0028] Step S64, for each of the to-be-verified set, the to-be-verified field information is respectively scored for credibility to obtain a corresponding credibility score, and the to-be-verified field information with the highest credibility score is output as the trusted field information;

[0029] Step S65, the trusted field information is generated as the field information of the preset demand field to generate the unique complete trusted data object associated with the corresponding business unique code, so as to enhance the object information value of the multi-source data.

[0030] Preferably, the step S64 comprises:

[0031] Step S641, for each of the to-be-verified field information in each of the to-be-verified set, the multi-source data as the data source of the to-be-verified field information is respectively scored for credibility, and the credibility level obtained according to the credibility level is configured as the credibility weight of the to-be-verified field information;

[0032] Step S642, for each of the to-be-verified set, the credibility score of each of the to-be-verified field information is calculated according to the respective credibility weight, and the to-be-verified field information with the highest credibility score is output as the trusted field information.

[0033] Preferably, in the step S642, the credibility score is calculated by using the following formula:

[0034]

[0035] Wherein, score is used to represent the credibility score, k is used to represent the number of credibility weights of the to-be-verified field information configuration, W i is used to represent each of the credibility weights, m is used to represent the number of multi-source data as the data source of the to-be-verified field information and associated with the corresponding credibility weight under each of the credibility levels, and w j is used to represent the preset weight of each of the multi-source data as the data source of the to-be-verified field information and associated with the corresponding credibility weight under each of the credibility levels.

[0036] Preferably, in the step S641:

[0037] When it is judged that the source of the multi-source data as the data source is single and reliable, the collection is verified, and there is a unique association with other information, the credibility level of the multi-source data is configured as strong association;

[0038] When it is judged that the source of the multi-source data as the data source is relatively reliable, the collection is natural collection, and there is a certain association with other information, the credibility level of the multi-source data is configured as general association;

[0039] When it is judged that the source of the multi-source data as the data source is temporarily reliable, the collection is natural collection, and there is a unique association with other information, the credibility level of the multi-source data is configured as weak association;

[0040] The credibility weight corresponding to the credibility level of strong association, the credibility weight corresponding to the credibility level of general association, and the credibility weight corresponding to the credibility level of weak association decrease in turn.

[0041] The application also provides a system for enhancing the value of object information based on multi-source data, which applies the method for enhancing the value of object information based on multi-source data.

[0042] A data acquisition module is configured to acquire multi-source data from each government affair system to form a data lake according to the big data analysis requirement of at least one business scene.

[0043] The first screening module is connected with the data acquisition module, and is configured to, for each business scenario, extract, from the data lake, the multi-source data satisfying a demand field missing standard and containing a corresponding business unique code as a standard data object according to a plurality of preset demand fields associated with the business scenario, and merge each standard data object with the same business unique code to form a corresponding standard object set;

[0044] The second screening module is connected with the first screening module, and is configured to, for each standard object set, obtain a corresponding target data object by screening each standard data object according to at least one business condition associated with the business scenario, and perform credibility identification on the target data object to screen out a corresponding credible data object, thereby obtaining a corresponding credible object set;

[0045] The information supplementing module is connected with the second screening module, and is configured to, for each credible object set, supplement a missing field information of each credible data object to obtain a corresponding supplemented data object containing all the preset demand fields;

[0046] The data matching module is connected with the information supplementing module, and is configured to match, according to the business unique code of each supplemented data object, the multi-source data associated with the business scenario and other than the standard data object from the data lake to obtain a matching data object;

[0047] The value enhancing module is connected with the information supplementing module and the data matching module, and is configured to, for each business unique code, process each supplemented data object and each matching data object to obtain a unique complete credible data object associated with the business unique code, thereby enhancing the object information value of the multi-source data.

[0048] The above technical solution has the following advantages or beneficial effects: based on a business scenario, multi-source data is used for data supplementing and credibility verification, the correlation between massive data is mined, the data island phenomenon is improved, the accuracy and completeness of data are effectively improved, the data has high information value, and the accuracy of government service data application is effectively improved, and the government-related work is efficiently carried out. BRIEF DESCRIPTION OF DRAWINGS

[0049] Figure 1 For a preferred embodiment of the present application, a flowchart of a method for enhancing object information value based on multi-source data is provided.

[0050] Figure 2 For a preferred embodiment of the present application, a sub-flowchart of step S2 is provided.

[0051] Figure 3 For the preferred embodiment of the present application, the sub-process schematic diagram of step S3 is shown in the following figure:

[0052] Figure 4 For the preferred embodiment of the present application, the sub-process schematic diagram of step S4 is shown in the following figure:

[0053] Figure 5 For the preferred embodiment of the present application, the sub-process schematic diagram of step S6 is shown in the following figure:

[0054] Figure 6 For the preferred embodiment of the present application, the sub-process schematic diagram of step S641 is shown in the following figure:

[0055] Figure 7 For the preferred embodiment of the present application, the structure schematic diagram of a system based on multi-source data enhanced object information value is shown in the following figure. DETAILED DESCRIPTION

[0056] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. The present application is not limited to this embodiment, and other embodiments can also fall within the scope of the present application as long as they comply with the main idea of the present application.

[0057] In the preferred embodiment of the present application, based on the above-mentioned problems existing in the prior art, a method for enhancing the value of object information based on multi-source data is provided, as shown in the following figure, which comprises: Figure 1

[0058] Step S1, according to the big data analysis requirements of at least one business scene, multi-source data is obtained from each government affair system to form a data lake;

[0059] Step S2, for each business scene, according to a plurality of preset requirement fields associated with the business scene, multi-source data satisfying a requirement field missing standard and containing a corresponding business unique code are extracted from the data lake as standard data objects, and each standard data object with the same business unique code is merged to form a corresponding standard object set;

[0060] Step S3, for each standard object set, according to at least one business condition of the preset associated business scene, each standard data object is screened to obtain a corresponding target data object, and the target data object is subjected to credibility identification to screen out a corresponding credible data object to obtain a corresponding credible object set;

[0061] Step S4, for each credible object set, each credible data object is subjected to missing field information supplement to obtain a corresponding supplemented data object containing all preset requirement fields;

[0062] ​Step S5, respectively according to the service unique code of each supplemented data object, the matching data object is obtained from the multi-source data in the data lake which matches the associated service scene and is in addition to the standard data object;

[0063] Step S6, for each service unique code, the unique complete trusted data object associated with the service unique code is processed according to each supplemented data object and each matching data object, so as to enhance the object information value of multi-source data.

[0064] Specifically, in the embodiment, when actually needing to perform big data analysis, it is possible to simultaneously analyze multiple service scenes, and then it is necessary to correspondingly acquire multi-source data under different service scenes. In order to simplify the data acquisition step, preferably, the big data components can be used to simultaneously access and converge the multi-source data of each service scene to form a data lake. When it is necessary to perform big data analysis on the data under one of the service scenes, it is only necessary to directly perform data extraction, without the need to perform the data acquisition step again.

[0065] It can be understood that, before performing big data analysis based on a certain service scene, the preset required fields that the data needs to contain are clear. Taking big data analysis of personal information as an example, if it is necessary to analyze the behavior characteristics of a type of crowd, such as the crowd in the same province and city, the preset required fields that the multi-source data at least needs to include are personal addresses. If it is necessary to further refine the crowd to a certain age group, the preset required fields that the multi-source data needs to include are ages. The data sources of personal information will also differ according to the business requirements, and the personal information recorded in the corresponding business system will also differ. For example, business system 1 needs to record ages, while business system 2 does not need to record ages. Therefore, the personal information from business system 2 lacks the preset required field of age, which causes information loss and directly affects the big data analysis result. Therefore, before performing big data analysis, it is necessary to enhance the object information value. Preferably, taking the personal information service scene as an example, the above-mentioned service unique code can be an identity card information. Taking the device information service scene as an example, the above-mentioned service unique code can be a device code, as long as it can represent the identity of the multi-source data description object.

[0066] Since the multi-source data is derived from different government systems, the data structures are various. Before performing data extraction from the data lake, preferably, the pre-processing of each multi-source data in the data lake is further included. The pre-processing mode includes but is not limited to de-privatization, cleaning and standardization, filtering of useless data, unification of data dictionary specification and name specification, etc.

[0067] After the above pre-processing, when performing data extraction from the data lake, preferably, the standard data object that can be used under the service scene is extracted through the preset required field and the set required field loss standard. The specific process is as follows Figure 2As shown, step S2 includes:

[0068] Step S21, respectively, statistics of each multi-source data in the missing service scene associated with each pre-set demand field missing field number of field;

[0069] Step S22, extract the missing field number of each multi-source data as the multi-source data that meets the demand field missing standard;

[0070] Step S23, extract each multi-source data containing business unique code from each multi-source data that meets the demand field missing standard as a standard data object, and merge each standard data object with the same business unique code to form a corresponding standard object set.

[0071] Specifically, in the embodiment, taking four preset demand fields as an example, i.e., field 1, field 2, field 3 and field 4, the multi-source data of the business scene in the data lake may only miss field 1, or miss field 1 and field 3, etc. As known, the more fields missing in the data, the more fields need to be supplemented, and the greater the corresponding supplement difficulty. In order to supplement as few field information as possible, it is preferred to extract each multi-source data with the least missing field as the multi-source data that meets the demand field missing standard. For example, for the four preset demand fields, each multi-source data missing only one field can be used as a standard data object. The missing fields in the standard data object are usually different, such as one multi-source data missing field 1, another missing field 2, and so on. Therefore, all multi-source data missing only one field are used as standard data objects.

[0072] In a preferred embodiment of the application, before step S23 is performed, it further includes filtering each multi-source data that meets the demand field missing standard according to a preset time window;

[0073] Then in step S23, the standard data object is extracted from each multi-source data that meets the demand field missing standard and the time window.

[0074] Specifically, in the embodiment, if only recent data trends need to be analyzed during big data analysis, multi-source data of other time windows does not meet the analysis requirements and needs to be filtered to ensure that multi-source data meets the needs of recent business and ensures the availability of data from the time dimension. The time window can be customized according to actual big data analysis requirements.

[0075] After the preliminary extraction of the standard data object, the standard data object here contains the preset demand field and meets the demand field missing standard and contains the corresponding business unique code, but not all of them are the data required by big data analysis. For example, only the behavior characteristics of individuals in Shanghai are needed to be analyzed, only the personnel information in Shanghai is needed. When extracting the standard data object, the preset demand field corresponding to this part is the address field. Therefore, it is necessary to further judge whether the personnel is an in-Shanghai object based on the field information of the address field. Therefore, the standard data object needs to be further screened. The specific process is as shown in Figure 3 Step S3 includes:

[0076] Step S31, for each standard object set, respectively, the field information of the preset demand field contained in each standard data object is analyzed, and when any characteristic analysis result indicates that the associated business condition is met, the corresponding standard data object is reserved as a target data object.

[0077] Step S32, respectively, the field information of the preset demand field contained in each target data object is standardized to obtain standardized field information, and according to a preset standardization degree, the credibility of each target data object is identified, so that the target data object corresponding to the standardized field information meeting the preset standardization degree is reserved as a credible data object, and a corresponding credible object set is obtained.

[0078] Specifically, in the embodiment, taking the need to screen in-Shanghai personnel as an example, the above-mentioned business condition is the recent behavior characteristics in Shanghai. If the business condition is met, the personnel is the in-Shanghai object needed, and the corresponding standard data object is the target data object. Taking the address field as an example, when personnel register address information, some business systems may have higher requirements for the detail of the address, and some business systems may have lower requirements for the detail of the address. Therefore, the address field obtained may only contain the community name. Different cities have a high probability of having the same community name. Therefore, the target data object whose field information of the address field only contains the community name is not a credible data object. However, the address detail is not that high, but it can be uniquely identified on the map, and it can be clear that the address belongs to Shanghai. Therefore, the corresponding target data object can be a credible data object, and the corresponding address field information can be temporarily considered as a credible address.

[0079] Preferably, when the standardization processing is performed, taking the address field information as an example, a neural network model trained in advance is preferably used to perform address field information extraction and entity name extraction, and then standardization processing is performed. The standardization rule can be set according to the address writing rule.

[0080] After screening the reliable data objects, for each reliable object set, the address registered by the person in one business system can contain province, city and street, and the address registered by the person in another business system can contain city and community, based on which, the registered addresses in the same reliable object set can be complemented to each other to perfect the address information, and the specific process is as shown in Figure 4 Step S4 includes:

[0081] Step S41, for each reliable object set, respectively count the missing preset demand fields of each reliable data object;

[0082] Step S42, for each missing preset demand field, respectively extract the field information of each reliable data object containing the preset demand field and perform deduplication processing;

[0083] Step S43, according to the field information after deduplication processing, respectively supplement the reliable data objects missing the preset demand field to obtain the corresponding supplemented data objects containing all preset demand fields.

[0084] Specifically, in this embodiment, taking the missing preset demand field as the address field as an example, after extracting the address field information of each reliable data object, the repeated address field information is deduplicated and then supplemented, and the unique address field information is directly supplemented, so that the address of the person registered in different business systems and having high reliability can be obtained, but considering the influence of data source reliability and many other factors, further reliability evaluation is needed to obtain a unique, accurate and complete address, and the specific process is as shown in Figure 5 Step S6 includes:

[0085] Step S61, for each preset demand field supplemented in each supplemented data object associated with each business unique code, respectively calculate the similarity between the supplemented missing field information in each supplemented data object and the corresponding field information of the matching data object;

[0086] Step S62, respectively extract the field information of the matching data object corresponding to the similarity greater than a threshold, and merge and deduplicate the extracted field information to obtain the merged and deduplicated field information associated with each preset demand field;

[0087] Step S63, add the supplemented missing field information and the merged and deduplicated field information associated with the same business unique code and preset demand field to a to-be-verified set as to-be-verified field information;

[0088] Step S64, for each to-be-verified set, respectively score the reliability of each to-be-verified field information to obtain a corresponding reliability score, and output the to-be-verified field information with the highest reliability score as the reliable field information;

[0089] Step S65, the trusted field information is generated as the field information of the preset demand field to generate the corresponding unique complete trusted data object associated with the unique service code, so as to enhance the object information value of the multi-source data.

[0090] Specifically, in the embodiment, since the multi-source data with the least missing fields is extracted as the multi-source data meeting the demand field missing standard, although the missing is the least, it is not necessarily trusted, and after a series of screening and supplementing processes based on the extracted part of the data, the field information with relatively high trustworthiness is obtained, but further supplement and trustworthiness evaluation are needed in combination with the multi-source data to further improve the data integrity and trustworthiness.

[0091] Preferably, when calculating the similarity, the pre-trained neural network model can also be used for word segmentation, and then the similarity is calculated. Based on the similarity, the similar addresses are merged and de-duplicated, and the trusted field information obtained has further improved trustworthiness. Further, the trustworthiness is evaluated in combination with the reliability and relevance of the data source, and the unique, accurate and complete data object is obtained based on the evaluation result, so as to enhance the information value.

[0092] In the preferred embodiment of the present application, as shown in Figure 6 Step S64 includes:

[0093] Step S641, for each trusted field information in each trusted set, the trustworthiness of the multi-source data as the data source of the trusted field information is graded respectively, and the trustworthiness level obtained according to the trustworthiness grading is configured as the corresponding trustworthiness weight of the trusted field information;

[0094] Step S642, for each trusted set, the trustworthiness score of each trusted field information associated with each trustworthiness weight is calculated respectively, and the trusted field information with the highest trustworthiness score is output as the trusted field information.

[0095] In the preferred embodiment of the present application, in step S642, the trustworthiness score is calculated by the following formula:

[0096]

[0097] Wherein, score represents the trustworthiness score, k represents the number of trustworthiness weights configured for the trusted field information, W i represents each trustworthiness weight, and m represents the number of multi-source data as the data source of the trusted field information and associated with the corresponding trustworthiness weight under each trustworthiness level, w jpreset weights of each multi-source data for representing data sources as field information to be verified under each credibility level and associating corresponding credibility weights.

[0098] Specifically, in the embodiment, the preset weights are pre-configured by expert experience, and the expert can preferably configure the preset weights by indicators such as source data type, data volume, non-empty proportion, and availability.

[0099] In the preferred embodiment of the application, in step S641:

[0100] When it is judged that the source of the multi-source data as the data source is single and reliable, the collection is verified, and there is a unique association with other information, the credibility level of the multi-source data is configured as strong association;

[0101] When it is judged that the source of the multi-source data as the data source is relatively reliable, the collection is natural collection, and there is a certain association with other information, the credibility level of the multi-source data is configured as general association;

[0102] When it is judged that the source of the multi-source data as the data source is temporarily reliable, the collection is natural collection, and there is a unique association with other information, the credibility level of the multi-source data is configured as weak association;

[0103] The credibility weight corresponding to the strong association credibility level, the credibility weight corresponding to the general association credibility level, and the credibility weight corresponding to the weak association credibility level decrease in turn.

[0104] Specifically, in the embodiment, the association levels of the strong association, the general association, and the weak association decrease in turn, and it can be understood that the higher the association level is, the higher the credibility of the data is, and therefore the greater the corresponding configured credibility weight is.

[0105] The application also provides a system for enhancing object information value based on multi-source data, which applies the method for enhancing object information value based on multi-source data. Figure 7 As shown in the figure, the system comprises:

[0106] A data acquisition module 1 is configured to acquire multi-source data from each government affair system to form a data lake according to the big data analysis requirement of at least one business scenario;

[0107] A first screening module 2 is connected to the data acquisition module 1 and is configured to, for each business scenario, extract multi-source data satisfying a demand field missing standard and containing a corresponding business unique code from the data lake as a standard data object according to a plurality of preset demand fields associated with the business scenario, and merge each standard data object with the same business unique code to form a corresponding standard object set.

[0108] The second screening module 3 is connected with the first screening module 2, and is used for screening each standard data object according to at least one business condition of a preset associated business scenario to obtain a corresponding target data object, and performing credibility identification on the target data object to screen out a corresponding credible data object, thereby obtaining a corresponding credible object set;

[0109] The information supplementing module 4 is connected with the second screening module 3, and is used for supplementing missing field information of each credible data object to obtain a corresponding supplemented data object containing all preset required fields for each credible object set;

[0110] The data matching module 5 is connected with the information supplementing module 4, and is used for matching and associating, according to a business unique code of each supplemented data object, a plurality of source data in a data lake to obtain a matched data object, wherein the plurality of source data are associated with the business scenario and are in addition to the standard data object;

[0111] The value enhancing module 6 is connected with the information supplementing module 4 and the data matching module 5, and is used for processing, according to each supplemented data object and each matched data object, a unique complete credible data object associated with the business unique code for each business unique code, so as to enhance the object information value of the plurality of source data.

[0112] The above only describes the preferred embodiments of the present application, and does not limit the implementation and protection scope of the present application. For those skilled in the art, it should be realized that any equivalent replacement and obvious change made according to the content of the present application should be included in the protection scope of the present application.

Claims

1. A method for enhancing object information value based on multi-source data, characterized in that, The method comprises the following steps: Step S1, obtaining multi-source data from each government affair system according to the big data analysis requirement of at least one business scenario to form a data lake; Step S2, for each business scenario, extracting multi-source data that meets a requirement field missing standard and contains a corresponding business unique code from the data lake as a standard data object according to a plurality of preset requirement fields associated with the business scenario, and merging each standard data object with the same business unique code to form a corresponding standard object set; Step S3, for each standard object set, filtering each standard data object according to at least one business condition associated with the business scenario to obtain a corresponding target data object, and performing credibility identification on the target data object to filter out a corresponding credible data object to obtain a corresponding credible object set; Step S4, for each credible object set, supplementing each credible data object to obtain a corresponding supplemented data object containing all the preset requirement fields; Step S5, respectively according to the business unique code of each supplemented data object, matching the multi-source data associated with the business scenario and excluding the standard data object from the data lake to obtain a matching data object; Step S6, for each business unique code, processing each supplemented data object and each matching data object to obtain a unique complete credible data object associated with the business unique code to enhance the object information value of the multi-source data; Step S6 comprises: Step S61, for each preset requirement field supplemented in each supplemented data object associated with each business unique code, respectively calculating the similarity between the missing field information supplemented in each supplemented data object and the field information of the matching data object corresponding to each supplemented data object; Step S62, extracting the field information of the matching data object corresponding to the similarity greater than a threshold value, and merging and deduplicating the extracted field information to obtain merged and deduplicated field information associated with each preset requirement field; Step S63, adding each missing field information and each merged and deduplicated field information associated with the same business unique code and the preset requirement field to a to-be-verified set as to-be-verified field information; Step S64, for each to-be-verified set, respectively performing credibility scoring on each to-be-verified field information to obtain a corresponding credibility score, and outputting the to-be-verified field information with the highest credibility score as credible field information; Step S65, generating a unique complete credible data object associated with the corresponding business unique code by taking the credible field information as the field information of the preset requirement field to enhance the object information value of the multi-source data.

2. The method for enhancing object information value based on multi-source data according to claim 1, characterized in that, The step S2 comprises: Step S21, respectively counting the number of missing fields of each preset requirement field associated with the business scenario in each multi-source data; Step S22, extracting each of the multi-source data with the least number of missing fields as the multi-source data satisfying the demand field missing criterion; Step S23, extracting each of the multi-source data containing the business unique code as the standard data object from the multi-source data satisfying the demand field missing criterion, and merging each of the standard data objects with the same business unique code to form the corresponding standard object set.

3. The method for enhancing object information value based on multi-source data according to claim 2, characterized in that, Before performing the step S23, further comprising filtering the multi-source data satisfying the demand field missing criterion according to a preset time window; Then in the step S23, extracting the standard data object from the multi-source data satisfying the demand field missing criterion and the time window.

4. The method for enhancing object information value based on multi-source data according to claim 1, characterized in that, The step S3 comprises: Step S31, for each standard object set, respectively performing feature analysis on the field information of the preset demand field contained in each standard data object, and when any feature analysis result indicates that the associated business condition is met, retaining the corresponding standard data object as the target data object; Step S32, respectively performing standardization processing on the field information of the preset demand field contained in each target data object to obtain standardized field information, and performing credibility identification on each target data object according to a corresponding preset standardization degree, so as to retain the target data object corresponding to the standardized field information meeting the preset standardization degree as the credible data object, and obtain the corresponding credible object set.

5. The method for enhancing object information value based on multi-source data according to claim 1, characterized in that, The step S4 comprises: Step S41, for each credible object set, respectively counting the preset demand fields missing in each credible data object; Step S42, for each missing preset demand field, respectively extracting the field information of each credible data object containing the preset demand field and performing deduplication processing; Step S43, according to each field information after deduplication processing, respectively supplementing the credible data object missing the preset demand field to obtain the corresponding data object after supplementing containing all the preset demand fields.

6. The method for enhancing object information value based on multi-source data according to claim 1, characterized in that, The step S64 comprises: Step S641, for each to-be-verified field information in each to-be-verified set, respectively performing credibility classification on the multi-source data as the data source of the to-be-verified field information, and configuring a corresponding credibility weight for the to-be-verified field information according to the credibility level obtained by credibility classification; Step S642, for each to-be-verified set, respectively calculating the credibility score of each to-be-verified field information associated according to each credibility weight, and outputting the to-be-verified field information with the highest credibility score as the credible field information.

7. The method of claim 6, wherein, In the step S642, the credibility score is calculated by using the following formula: ; wherein, for representing the credibility score, for representing the number of credibility weights of the to-be-verified field information configuration, for representing each of the credibility weights, for representing the number of multi-source data as data sources of the to-be-verified field information under each of the credibility levels and associated with the corresponding credibility weights, for representing a preset weight of each of the multi-source data as data sources of the to-be-verified field information under each of the credibility levels and associated with the corresponding credibility weights. 8.The method of claim 6, wherein, In the step S641: When the source of the multi-source data as the data source is single and reliable, the collection is verified and has a unique association with other information, the credibility level of the multi-source data is configured as strong association; when it is judged that the source of the multi-source data as a data source is relatively reliable, collected as natural collection, not verified but has a certain correlation with other information, the credibility level of the multi-source data is configured as general correlation; when it is judged that the source of the multi-source data as a data source is temporarily reliable, collected as natural collection, not verified and has only correlation with other information, the credibility level of the multi-source data is configured as weak correlation; the credibility weight corresponding to the credibility level of strong correlation, the credibility weight corresponding to the credibility level of general correlation and the credibility weight corresponding to the credibility level of weak correlation decrease in turn. 9.A system for enhancing value of object information based on multi-source data, characterized in that, The system based on the method for enhancing the object information value of multi-source data according to any one of claims 1-8, the system comprises: a data acquisition module configured to acquire multi-source data from each government affair system to form a data lake according to the big data analysis requirement of at least one business scenario; a first screening module connected to the data acquisition module and configured to, for each business scenario, extract the multi-source data satisfying a demand field missing standard and containing a corresponding business unique code from the data lake as a standard data object according to a plurality of preset demand fields associated with the business scenario, and merge each standard data object with the same business unique code to form a corresponding standard object set; a second screening module connected to the first screening module and configured to, for each standard object set, screen each standard data object according to at least one business condition associated with the business scenario to obtain a corresponding target data object, and perform credibility identification on the target data object to screen out a corresponding credible data object to obtain a corresponding credible object set; an information supplementing module connected to the second screening module and configured to, for each credible object set, supplement the missing field information of each credible data object to obtain a corresponding supplemented data object containing all the preset demand fields; a data matching module connected to the information supplementing module and configured to match the multi-source data associated with the business scenario and other than the standard data object from the data lake according to the business unique code of each supplemented data object to obtain a matching data object; a value enhancing module connected to the information supplementing module and the data matching module and configured to, for each business unique code, process each supplemented data object and each matching data object to obtain a unique complete credible data object associated with the business unique code to enhance the object information value of the multi-source data.

Citation Information

Patent Citations

  • Data fusion method and device

    CN110119413A

  • Generating data pattern information

    US20120197887A1