Enterprise background investigation method and system
Through the tag attribute collection of database data, the source weight and correlation coefficient are set for data fusion, the problem of insufficient data integration accuracy in enterprise background investigations is solved, and a more accurate and safe back-click report generation is achieved.
Patent Information
- Application Number
- CN202510864775.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the process of data integration of the prior art enterprise background investigation platforms, in the process of data integration of data with the same label attributes, resulting in insufficient data accuracy and affecting the accuracy of back-tuning results.
Through the tag attributes, the data to be used in several databases are automatically collected, the source weight is set and data fusion is fusion based on the correlation coefficient, and the configuration module is built to generate back-tuning reports, realize a fusion strategy of dynamic switching, adapt to different data fusion scenarios, and ensure data security through access levels.
It improves the accuracy of data fusion, improves the accuracy of back-tuning results, and ensures data security, adapts to the generation of self-service customized reports that meet the needs of different users.
Smart Images

Figure CN120371885A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and system for investigating the background of an enterprise. Background Art
[0002] Background investigation is a process where an independent professional organization verifies and compares the background information of the subject of investigation based on authoritative data sources and forms an investigation report. When the subject of investigation is a company, the company background investigation helps to fully understand its operating conditions and potential risks through data analysis and information verification.
[0003] In terms of data processing, the corporate background investigation platforms currently on the market need to connect to multiple data sources to acquire data. During the data integration process, there is a lack of data fusion methods for data with the same label attributes, resulting in insufficient data precision and affecting the accuracy of background investigation results. Summary of the invention
[0004] In view of the deficiencies in the prior art, the purpose of the present invention is to provide a corporate background investigation method and system, aiming to solve the technical problem in the prior art that when conducting corporate background investigations and integrating data from multiple data sources, there is a lack of data fusion means for data with the same label attributes, resulting in insufficient data accuracy and affecting the accuracy of background check results.
[0005] In order to achieve the above objectives, in a first aspect, an embodiment of the present application provides a method for investigating an enterprise background, comprising the following steps: Acquire several label attributes of the background check target, and acquire a same-label data group corresponding to the label attributes from several databases based on the label attributes, wherein the same-label data group includes several stand-by data; In the same data group with the same label, a plurality of the unused data are respectively assigned source weights, and fusion data corresponding to the label attribute is obtained based on the source weights and the unused data; The tag attributes and the fused data are combined into category data, access levels are set for different category data, and correspond to different configuration modules, and a background check report is generated based on the user's generation requirements and the configuration module.
[0006] Furthermore, the step of acquiring a data group with the same label corresponding to the tag attribute from a plurality of databases based on the tag attribute, wherein the data group with the same label includes a plurality of unused data, comprises: Selecting a plurality of target databases from the plurality of databases based on the tag attributes; Acquire background data corresponding to the tag attribute from the target database, and pre-process a number of the background data to form a number of stand-by data; Correspond a number of the to-be-used data to a number of the label attributes respectively to obtain a number of same-label data groups.
[0007] Furthermore, the preprocessing includes filling-in-the-blank processing, abnormal replacement and standardization processing.
[0008] Furthermore, the step of respectively assigning source weights to a number of the to-be-used data and obtaining fusion data corresponding to the label attributes based on the source weights and the to-be-used data includes: Set source weights based on the credibility of the target database corresponding to the to-be-used data, and extract the feature values corresponding to the label attributes in a number of the to-be-used data; Select the feature value corresponding to the target database with the highest source weight as the reference feature value, and select the remaining feature values as reference feature values; Obtain a number of correlation coefficients between the reference feature value and a number of the reference feature values, and compare the number of the correlation coefficients with a coefficient threshold respectively to obtain fusion data.
[0009] Furthermore, the acquisition formula of the correlation coefficient is: , wherein, represents the correlation coefficient between the reference feature value and the i-th reference feature value, represents the reference feature value, represents the i-th reference feature value, represents the mean value between the reference feature value and the i-th reference feature value, and .
[0010] Furthermore, the step of comparing the number of the correlation coefficients with the coefficient threshold respectively to obtain fusion data includes: If all the number of the correlation coefficients are greater than the coefficient threshold, select the feature value corresponding to the target database with the highest source weight as the fusion data; If any one of the correlation coefficients is less than the coefficient threshold, fuse a number of the feature values based on the source weights into fusion data.
[0011] Furthermore, the step of generating a background check report based on the generation requirements of the user and the configuration module includes: Select to-be-used modules from a number of the configuration modules according to the generation requirements of the user; Obtain the permission level of the user, and compare the permission level with the access level of the to-be-used modules; Select the standby module corresponding to the access level lower than the permission level as the final module, and combine the category data corresponding to the final module into a background investigation report.
[0012] In a second aspect, an embodiment of the present application provides an enterprise background investigation system, which is applied to the enterprise background investigation method described in the first aspect above. The system includes: An extraction module, configured to obtain a plurality of tag attributes of a background investigation target, and obtain a same-tag data group corresponding to the tag attributes from a plurality of databases based on the tag attributes. The same-tag data group includes a plurality of standby data; A combination module, configured to assign source weights to a plurality of the standby data respectively within the same same-tag data group, and obtain fusion data corresponding to the tag attributes based on the source weights and the standby data; An execution module, configured to combine the tag attributes and the fusion data into category data, set access levels for different category data, and correspond them to different configuration modules, and generate a background investigation report based on the generation requirements of the user and the configuration modules.
[0013] In a third aspect, an embodiment of the present application provides a computer, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the enterprise background investigation method described in the first aspect above is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the enterprise background investigation method described in the first aspect above is implemented.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the tag attributes, the standby data related to the tag attributes in a plurality of databases are automatically collected, and the preliminary integration of the same type of data is completed; by setting the source weights, the credibility of the databases is used as the basis for fusion. When the correlation coefficient is relatively high, the data of the database with the highest credibility is directly adopted to avoid data conflicts. When the correlation coefficient is relatively low, weighted fusion takes into account the credibility of multiple sources, improves the accuracy of data fusion, provides an effective and accurate fusion means for data with the same tag attributes, improves the accuracy of the background investigation results, adopts two fusion methods to form a dynamically switched fusion strategy, and adapts to different data fusion scenarios; by constructing the configuration module, the background report can be customized according to user needs; by setting the access levels, the category data is classified by permission to avoid data leakage and ensure data security. Description of the Drawings
[0016] Figure 1It is a flow chart of the enterprise background investigation method in the first embodiment of the present invention; Figure 2 It is a structural block diagram of the enterprise background investigation system in the second embodiment of the present invention; The following specific implementation manner will further illustrate the present invention in conjunction with the above-mentioned drawings. DETAILED DESCRIPTION
[0017] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0018] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0020] See also Figure 1 The enterprise background investigation method provided by the first embodiment of the present invention comprises the following steps: S10: Acquire several label attributes of the background check target, and acquire a same-label data group corresponding to the label attributes from several databases based on the label attributes, wherein the same-label data group includes several stand-by data; The step S10 comprises: S110: selecting a plurality of target databases from the plurality of databases based on the tag attributes; Taking Company A as the background check target as an example, one of the label attributes is registered capital, then the registered capital of Company A is searched from several databases, and the database with corresponding data is selected as the target database.
[0021] S120: Acquire background data corresponding to the tag attribute from the target database, and pre-process a plurality of the background data to form a plurality of stand-by data; The preprocessing includes gap filling, anomaly replacement and standardization. The gap filling refers to the supplement of the data after the data is missing. For example, if a certain label attribute is registered capital, and there is part of the missing data of registered capital in a database, it can be predicted and supplemented through industry average or random forest regression method. The anomaly replacement refers to the replacement of abnormal data values. For example, if a certain label attribute is the number of employees, and the number of employees in a database is negative, it will be replaced with a positive value. The standardization refers to the unified formatting of the data. For example, if a certain label attribute is registered capital, the registered capital in one database is 5,000,000 yuan, and the registered capital in another database is 5 million yuan, then 5 million yuan will also be converted into 5,000,000 yuan.
[0022] S130: respectively correspond the plurality of unused data to the plurality of label attributes to obtain a plurality of data groups with the same label; Still taking the tag attribute as the registered capital, three pieces of standby data are extracted from three databases respectively, and then the three pieces of standby data are classified into a same-label data group.
[0023] S20: in the same data group with the same label, assigning source weights to a plurality of the unused data respectively, and acquiring fused data corresponding to the label attribute based on the source weights and the unused data; The step S20 comprises: S210: setting a source weight based on the credibility of the target database corresponding to the stand-by data, and extracting a plurality of feature values corresponding to the label attribute in the stand-by data; Credibility represents the authority of the target database. The higher the credibility, the greater the value of the source weight of the target database. It should be noted that the sum of the source weights of the target database corresponding to the same stand-by data is 1.
[0024] S220: selecting the feature value corresponding to the target database with the highest source weight as the benchmark feature value, and selecting the remaining feature values as reference feature values; S230: obtaining a plurality of correlation coefficients between the benchmark feature value and a plurality of the reference feature values, and comparing the plurality of correlation coefficients with coefficient thresholds respectively to obtain fused data; The formula for obtaining the correlation coefficient is: , in, represents the correlation coefficient between the benchmark eigenvalue and the i-th reference eigenvalue, represents the benchmark eigenvalue, represents the i-th reference eigenvalue, represents the mean between the benchmark eigenvalue and the i-th reference eigenvalue, and ; Furthermore, if several of the correlation coefficients are greater than the coefficient threshold, the feature value corresponding to the target database with the highest source weight is selected as the fused data; if any of the correlation coefficients is less than the coefficient threshold, several of the feature values are weighted and fused into the fused data based on the source weight. In this embodiment, the coefficient threshold is 0.9.
[0025] S30: combining the tag attribute and the fused data into category data, setting access levels for different category data, and corresponding them to different configuration modules, and generating a background check report based on the user's generation requirements and the configuration module; It can be understood that by selecting the configuration module, the category data corresponding to the configuration module can be extracted. By selecting different configuration modules, different category data can be combined into the background check report. Specifically, step 30 includes: S310: selecting a standby module from a plurality of configuration modules according to a user's generation requirement; The generation requirement refers to the tag attributes related to the background check target required by the user, and several configuration modules can be indexed through several tag attributes.
[0026] S320: Obtaining the user's authority level, and comparing the authority level with the access level of the standby module; S330: Select the standby module corresponding to the access level lower than the authority level as the final module, and combine the category data corresponding to the final module into a background check report.
[0027] Furthermore, if the authority level is lower than the access level, an authorization instruction is triggered, requiring the user to provide an authorization file corresponding to the access level of the standby module to complete data acquisition.
[0028] The tag attributes are used to automatically collect the unused data related to the tag attributes in several databases to complete the preliminary integration of similar data; by setting the source weight, the credibility of the database is used as the basis for fusion. When the correlation coefficient is high, the data of the database with the highest credibility is directly used to avoid data conflicts. When the correlation coefficient is low, weighted fusion takes into account the credibility of multiple sources to improve the accuracy of data fusion. Two fusion methods are used to form a dynamically switched fusion strategy to adapt to different data fusion scenarios; by constructing the configuration module, the background report can be customized according to user needs; by setting the access level, the category data is graded for permissions to avoid data leakage and ensure data security.
[0029] See also Figure 2 The second embodiment of the present invention provides a corporate background investigation system, which is applied to the corporate background investigation method described in the above embodiment. The description has been made and will not be repeated. As used below, the terms "module", "unit", "subunit" and the like can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0030] The system comprises: The extraction module 10 is used to obtain several label attributes of the background check target, and based on the label attributes, obtain a same-label data group corresponding to the label attributes from several databases, wherein the same-label data group includes several stand-by data; The extraction module 10 comprises: A first unit is used to select a plurality of target databases from a plurality of the databases based on the tag attribute; The second unit is used to obtain background data corresponding to the tag attribute from the target database, and pre-process a plurality of the background data to form a plurality of stand-by data; The third unit is used to correspond the plurality of unused data to the plurality of label attributes respectively, so as to obtain a plurality of data groups with the same label; A combining module 20 is used to assign source weights to a plurality of the unused data in the same data group with the same label, and obtain fused data corresponding to the label attribute based on the source weights and the unused data; The combined module 20 comprises: A fourth unit is used to set a source weight based on the credibility of the target database corresponding to the stand-by data, and extract a plurality of feature values corresponding to the label attribute in the stand-by data; A fifth unit, configured to select a feature value corresponding to the target database having the highest source weight as a benchmark feature value, and select the remaining feature values as reference feature values; A sixth unit is used to obtain a plurality of correlation coefficients between the benchmark feature value and a plurality of the reference feature values, and compare the plurality of correlation coefficients with coefficient thresholds respectively to obtain fused data; The sixth unit is specifically used for selecting the feature value corresponding to the target database with the highest source weight as the fused data if several of the correlation coefficients are greater than the coefficient threshold; if any of the correlation coefficients is less than the coefficient threshold, weighting and fusing several of the feature values into the fused data based on the source weight; An execution module 30, for combining the tag attributes and the fused data into category data, setting access levels for different category data, and corresponding to different configuration modules, and generating a background check report based on the user's generation requirements and the configuration module; The execution module 30 includes: The seventh unit is used to select a standby module from a plurality of configuration modules according to the user's generation requirements; An eighth unit is used to obtain the user's authority level and compare the authority level with the access level of the standby module; The ninth unit is used to select the standby module corresponding to the access level lower than the authority level as the final module, and combine the category data corresponding to the final module into a background check report.
[0031] The present invention also provides a computer, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the enterprise background investigation method as described in the above technical solution when executing the computer program.
[0032] The present invention also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the enterprise background investigation method described in the above technical solution is implemented.
[0033] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0034] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. An enterprise background investigation method, characterized in that, The following steps are involved: Acquire several label attributes of the background check target, and acquire a same-label data group corresponding to the label attributes from several databases based on the label attributes, wherein the same-label data group includes several stand-by data; In the same data group with the same label, a plurality of the unused data are respectively assigned source weights, and fusion data corresponding to the label attribute is obtained based on the source weights and the unused data; The tag attributes and the fused data are combined into category data, access levels are set for different category data, and correspond to different configuration modules, and a background check report is generated based on the user's generation requirements and the configuration module.
2. The enterprise background investigation method according to claim 1, characterized in that The step of acquiring a data group with the same label corresponding to the tag attribute from a plurality of databases based on the tag attribute, wherein the data group with the same label includes a plurality of unused data comprises: Selecting a plurality of target databases from the plurality of databases based on the tag attributes; Acquire background data corresponding to the tag attribute from the target database, and pre-process a number of the background data to form a number of stand-by data; A plurality of the unused data are respectively mapped to a plurality of the label attributes to obtain a plurality of data groups with the same label.
3. The enterprise background investigation method according to claim 2, wherein The preprocessing includes gap filling, abnormal replacement and standardization.
4. The enterprise background investigation method according to claim 2, wherein The step of assigning source weights to the plurality of unused data respectively and acquiring fused data corresponding to the label attribute based on the source weights and the unused data comprises: Setting a source weight based on the credibility of a target database corresponding to the standby data, and extracting a plurality of feature values corresponding to the label attribute in the standby data; Selecting the eigenvalue corresponding to the target database with the highest source weight as the benchmark eigenvalue, and selecting the remaining eigenvalues as reference eigenvalues; A plurality of correlation coefficients between the benchmark eigenvalue and a plurality of the reference eigenvalues are obtained, and the plurality of correlation coefficients are respectively compared with coefficient thresholds to obtain fused data.
5. The enterprise background investigation method according to claim 4, characterized in that, The formula for obtaining the correlation coefficient is: , Among them, represents the correlation coefficient between the reference feature value and the i-th reference feature value, represents the reference feature value, represents the i-th reference feature value, represents the mean value between the reference feature value and the i-th reference feature value, and .
6. The enterprise background investigation method according to claim 4, wherein The step of comparing the plurality of correlation coefficients with the coefficient thresholds respectively to obtain fused data comprises: If a number of the correlation coefficients are all greater than the coefficient threshold, the feature value corresponding to the target database with the highest source weight is selected as the fused data; If any of the correlation coefficients is less than the coefficient threshold, a plurality of the feature values are weighted and fused into fused data based on the source weights.
7. The enterprise background investigation method according to claim 1, characterized in that, The step of generating a background check report based on the user's generation requirements and the configuration module includes: Selecting a standby module from a plurality of configuration modules according to the user's generation requirements; Obtaining the user's authority level, and comparing the authority level with the access level of the stand-by module; The standby module corresponding to the access level lower than the authority level is selected as the final module, and the category data corresponding to the final module is combined into a background check report.
8. An enterprise background investigation system, which is applied to the enterprise background investigation method described in any one of claims 1 to 7, and is characterized in that, The system comprises: An extraction module, used to obtain several label attributes of the background check target, and based on the label attributes, obtain a same-label data group corresponding to the label attributes from several databases, wherein the same-label data group includes several stand-by data; A combination module, used for assigning source weights to a plurality of the unused data in the same data group with the same label, and obtaining fused data corresponding to the label attribute based on the source weights and the unused data; An execution module is used to combine the tag attributes and the fused data into category data, set access levels for different category data, and correspond them to different configuration modules, and generate a background check report based on the user's generation requirements and the configuration module.
9. A computer, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the enterprise background investigation method as described in any one of claims 1 to 7 is implemented.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the enterprise background investigation method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Label-based data analysis method and analysis system
CN109062984A
Industrial capability labeling method based on big data technology
CN113641864A
Enterprise credit data processing method and device
CN114138869A
Data sharing method with visual authority
CN117592113A
Enterprise credit survey data processing system and method based on multi-dimensional data
CN118628128A