Data matching method and device, medium and electronic equipment

By combining multi-round progressive data matching processing with data source credibility levels, the problems of mismatch and missed match caused by inconsistent data formats and differences in identification rules in the credit reporting platform have been solved, thereby improving the accuracy of data integration and the reliability of credit assessment.

CN121786504AActive Publication Date: 2026-04-03QIANTANG CREDIT INFORMATION CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-03
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In credit reporting platforms, inconsistent data formats and different data identification rules from different data sources lead to mismatches and missed matches, affecting the accuracy of data integration and the reliability of credit assessment.

Method used

By using a multi-round progressive data matching process, combined with the credibility level of the data source, we ensure that high-credibility data plays a key role in subsequent matching, reducing the risk of mismatches. Furthermore, by incorporating new data with higher credibility in each round of matching, we form an iterative and optimized matching process that effectively links scattered data of the same subject.

Benefits of technology

It significantly improves the accuracy and completeness of personal data on the credit reporting platform, reduces the risk of mismatch, and enhances the reliability of credit assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121786504A_ABST
    Figure CN121786504A_ABST
Patent Text Reader

Abstract

The invention provides a data matching method and device, a medium and electronic equipment, and in the method, each piece of to-be-processed data sent by each data source can be received, at least two rounds of data matching are performed on each piece of received to-be-processed data, and each matching result set is obtained for a downstream service system to call. In any round of data matching, target data required in the round of data matching can be determined from the to-be-processed data according to a preset credibility level corresponding to each data source, and the target data and matched data contained in a matching result set output through the previous round of data matching serve as input data. And inputting each piece of input data into a preset matching model, so that the matching model identifies the input data belonging to the same entity object from each piece of input data and merges the input data into the same matching result set to obtain a matching result set output after the round of data matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and more particularly to a data matching method, apparatus, medium, and electronic device. Background Technology

[0002] Currently, credit reporting platforms can collect users' personal data through various channels (such as repayment records, personal loan repayment records, credit inquiry records, and consumption records). After matching and processing data from different data sources belonging to the same entity, these platforms can, on the one hand, provide banks, consumer finance companies, and other enterprises with comprehensive user credit profiles and credit reports, enabling them to accurately assess credit risk and ensure business compliance; on the other hand, they can provide individual users with objective credit assessment criteria, directly impacting their access to credit interest rates and related service rights. Therefore, the completeness and accuracy of personal data within credit reporting platforms are the fundamental guarantee for the reliability and application value of their credit assessment results.

[0003] However, during the matching process between internal data and external multi-source data on the credit reporting platform, two core problems can easily arise due to inconsistent data formats and data identification rules among different data sources (e.g., different anonymization formats for ID numbers, inconsistent name entry standards, etc.): First, mismatch, which incorrectly associates data from different entities with similar characteristics; second, missed match, which fails to effectively associate valid data of the same entity. These two problems directly affect the accuracy of data integration, and thus the reliability of credit assessment. Summary of the Invention

[0004] In view of the above, one or more embodiments of this specification provide the following technical solutions: According to a first aspect of one or more embodiments of this specification, a data matching method is proposed, the method comprising: Receive data to be processed from various data sources; At least two rounds of data matching are performed on each piece of data to be processed to obtain each matching result set for use by downstream business systems; wherein, each matching result set is used to represent a data profile of an entity object; In any round of data matching: Based on the preset credibility level of each data source, the target data required for this round of data matching is determined from each data to be processed. The higher the round of data matching, the higher the credibility level of the data source corresponding to the determined target data. Each target data and each matched data included in the matching result set output from the previous round of data matching are used as input data. Each input data is then input into a preset matching model so that the matching model can identify input data belonging to the same entity object from each input data and merge them into the same matching result set to obtain the matching result set output from this round of data matching.

[0005] According to a second aspect of one or more embodiments of this specification, a business execution method is provided, the method comprising: Receive a data retrieval request; the data retrieval request contains filtering conditions, which include at least one of the following: specifying a credibility level, specifying a matching model type, and specifying a business scenario. The target matching result set that matches the filtering criteria is determined from each pre-processed matching result set and returned.

[0006] According to a third aspect of one or more embodiments of this specification, a data matching apparatus is provided, comprising: The receiving module is used to receive the data to be processed sent by each data source; The matching module is used to perform at least two rounds of data matching on each piece of data to be processed to obtain matching result sets for use by downstream business systems. Each matching result set is used to represent a data profile of an entity object. In any round of data matching: based on the preset credibility level corresponding to each data source, the target data required for this round of data matching is determined from each piece of data to be processed. The higher the round of data matching, the higher the credibility level of the data source corresponding to the determined target data. The target data and the matched data contained in the matching result set output by the previous round of data matching are used as input data. The input data are input into a preset matching model so that the matching model can identify input data belonging to the same entity object from the input data and merge them into the same matching result set to obtain the matching result set output by this round of data matching.

[0007] According to a fourth aspect of one or more embodiments of this specification, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor performs the executable instructions to implement the steps of the data matching method described above.

[0008] According to a fifth aspect of one or more embodiments of this specification, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the steps of the data matching method described above.

[0009] According to a sixth aspect of one or more embodiments of this specification, a computer program product is provided, comprising a computer program / instructions that, when executed by a processor, implement the steps of the data matching method described above.

[0010] As can be seen from the above embodiments, the system first receives the data to be processed from each data source, performs at least two rounds of data matching on the received data to obtain matching result sets for use by downstream business systems. Each matching result set is used to represent the data profile of an entity object. In any round of data matching: based on the preset credibility level corresponding to each data source, the target data required for this round of data matching is determined from the data to be processed. The higher the round of data matching, the higher the credibility level of the data source corresponding to the determined target data. The target data and the matched data contained in the matching result set output by the previous round of data matching are used as input data. The input data are input into a preset matching model so that the matching model can identify input data belonging to the same entity object from the input data and merge them into the same matching result set to obtain the matching result set output by this round of data matching.

[0011] This method, on the one hand, increases the credibility level of the data source corresponding to the target data in each round, ensuring that high-credibility data plays a key calibration role in subsequent matching and reducing the risk of mismatches caused by inconsistent formats and different labeling rules of low-credibility data. On the other hand, each round of matching is based on the results of the previous round and incorporates new data with higher credibility, effectively linking scattered data of the same subject and reducing the problem of missed matches caused by non-standard data labeling. Ultimately, the entity object data profile formed through multiple rounds of matching significantly improves the accuracy of personal data on the credit reporting platform. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the architecture of a data matching system provided in an exemplary embodiment.

[0013] Figure 2 This is a flowchart illustrating a data matching method provided in an exemplary embodiment.

[0014] Figure 3 This is a schematic diagram of any round of data matching provided in an exemplary embodiment.

[0015] Figure 4 This is a schematic diagram of multi-round data matching provided in an exemplary embodiment.

[0016] Figure 5 This is a flowchart illustrating a business execution method provided in an exemplary embodiment.

[0017] Figure 6 This is a schematic structural diagram of a device provided in an exemplary embodiment.

[0018] Figure 7 This is a block diagram of a data matching device provided in an exemplary embodiment.

[0019] Figure 8 This is a block diagram of a data matching device provided in an exemplary embodiment. Detailed Implementation

[0020] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0021] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.

[0022] Currently, credit reporting platforms can collect users' personal data from various data sources, clean, match, and process it to form a standardized personal credit dataset. Then, when a user requests to execute a corresponding business from any business platform, after obtaining the user's authorization, the business platform can retrieve the user's personal credit data from the credit reporting platform to assess the user's credit risk.

[0023] For example, when a user requests a deposit-free rental service on a housing rental platform, the platform can send a request to the credit reporting platform to obtain the user's personal credit data after obtaining the user's authorization. After verifying the validity of the authorization, the credit reporting platform can respond to the received request and return the user's standardized personal credit dataset (such as historical repayment records, credit usage, default information, etc.) to the housing rental platform. Based on this, the housing rental platform can determine the user's credit risk based on the obtained personal credit dataset and decide whether to accept the user's deposit-free rental application.

[0024] For example, when a user applies for "buy now, pay later" or "credit purchase" services on an e-commerce platform, the platform can, with the user's authorization, request a personal credit data query from a credit reporting agency. After verifying the validity of the authorization, the credit reporting agency returns the user's standardized personal credit dataset to the e-commerce platform. This dataset allows the platform to assess the user's willingness and ability to fulfill their obligations, and ultimately decide whether to grant them a credit payment limit and corresponding interest-free period or installment terms.

[0025] However, during the matching process between internal data of the credit reporting platform (i.e., the personal credit data of some users already held by the credit reporting platform) and external multi-source data, two core problems are prone to occur due to the inconsistent data formats and data identification rules of different data sources (e.g., different anonymization formats for ID numbers, inconsistent name entry standards, etc.): First, mismatch, that is, incorrectly associating data from different entities with similar characteristics; second, missed match, that is, failing to effectively associate valid data of the same entity. These two types of problems directly affect the accuracy of data integration, and thus affect the reliability of credit assessment.

[0026] Based on this, this specification provides a data matching method that achieves accurate data integration through multi-round progressive data matching processing combined with the credibility level of the data source. On the one hand, by increasing the credibility level of the data source corresponding to the target data in each round, it ensures that high-credibility data plays a key calibration role in subsequent matching, reducing the risk of mismatches caused by inconsistent formats and differences in identification rules of low-credibility data. On the other hand, each round of matching is based on the results of the previous round and incorporates new data with higher credibility, forming a continuously iterative and optimized matching process. This effectively associates scattered data of the same subject, reducing the problem of missed matches caused by non-standard data identification. Finally, the entity object data profile formed through multi-round matching significantly improves the completeness and accuracy of personal data on the credit reporting platform. The technical solution of this invention will be described in detail below with reference to specific embodiments.

[0027] Figure 1 This is a schematic diagram of the architecture of a data matching system provided in an exemplary embodiment. For example... Figure 1 As shown, the system may include a client 11, a credit reporting platform server 12, a business platform server 13, a network 14, etc.

[0028] The client 11 can be a user's electronic device, used to initiate requests or operation commands and interact with the credit reporting platform's server 12 or the business platform's server 13. These electronic devices can be, but are not limited to: personal computers (PCs), mobile phones, tablets, laptops, PDAs, and wearable devices (such as smart glasses, smartwatches, etc.).

[0029] During operation, client 11 can run a client-side program of an application to implement the relevant functions of that application. For example, when the client runs a housing rental service application, it can function as a corresponding housing rental service client. This housing rental service client application can be launched and run on an electronic device. The client-side program can be a native application installed on the electronic device, or it can be a mini-program, quick app, or other similar form. When using web technologies such as HTML5, the relevant functions can be implemented through a page displayed in a browser. This browser can be a standalone browser application or a browser module embedded in some applications.

[0030] The server 12 of the credit reporting platform can be a physical server containing an independent host, or a virtual server hosted in a host cluster. It is used to receive data requests from the client 11 or the business platform server 13, and respond to these requests to provide credit assessment-related services.

[0031] Specifically, the server-side 12 of the credit reporting platform can perform standardized processing, quality verification, and multi-round matching and fusion operations on the received user data, thereby constructing a matching result set corresponding to each entity object. This set is used to represent the credit profile or other relevant data profile of the entity object, which can be queried by users or called by downstream business systems. In addition, the server-side 12 of the credit reporting platform is also responsible for ensuring data security and privacy protection, ensuring that all operations are carried out with the user's authorization and in compliance with relevant laws and regulations.

[0032] The server 13 of the business platform can be a core component of a specific type of business application system (e.g., the business application system provided by a housing rental platform or the business application system provided by a bank loan approval platform). It is used to interact with the server 12 and client 11 of the credit reporting platform, and obtain the user's personal credit data by calling the services provided by the credit reporting platform in order to make more accurate business decisions.

[0033] For example, in a housing rental scenario, the business platform's server 13 can send a query request to the credit reporting platform's server 12 after receiving a user's application for a deposit-free rental, and decide whether to approve the user's application based on the returned credit assessment results.

[0034] The business platform's server 13 can also be deployed on a physical server or a virtual server to run programs with specific business logic, supporting access from various forms of clients, such as native applications and web applications.

[0035] As for the network 14 that facilitates interaction between the client 11, the credit reporting platform server 12, and the business platform server 13, communication can be achieved using wired or wireless networks based on the communication methods supported by the corresponding electronic devices. This manual does not impose any restrictions on this. For example, when the client 11 runs on a PC, it supports wired / wireless networks. When the client 11 runs on mobile terminals such as mobile phones or wearable devices, it mainly uses wireless networks such as 4G / 5G / Wi-Fi. To ensure data security, encrypted wired networks or dedicated wireless networks can be used for transmission.

[0036] Figure 2 This is a flowchart illustrating a data matching method provided in an exemplary embodiment, including: S200: Receives data to be processed from various data sources.

[0037] In this specification, the server of the credit reporting platform can receive data to be processed from at least one preset data source, and perform standardization processing, quality verification, and subsequent multi-round matching and fusion operations on the data to be processed, and finally construct the matching result set corresponding to each entity object for downstream business systems to call.

[0038] In the above context, different data sources can refer to various applications, platforms, or systems that are authorized by users, comply with personal information protection and credit reporting business management requirements, and are able to provide credit dimension data or feature data related to entities (such as individuals, enterprises, regions, etc.).

[0039] The aforementioned data to be processed can refer to multi-dimensional raw data sent from the different data sources mentioned above, such as: facial recognition data collected from the identity verification business platform, travel data from the railway ticketing platform, social security data from the government information disclosure platform, travel preference data from the travel service platform, check-in behavior data from the accommodation service platform, remote area label data from the civil affairs service platform, mobile phone model from the mobile device operating system, mobile phone login logs, etc.

[0040] Furthermore, the aforementioned data to be processed can also refer to the user's own data uploaded by the user through the self-upload portal provided by the credit reporting platform's application, mini-program, or web interface. For example, users can proactively upload their personal asset certificates (such as scanned copies of property ownership certificates and vehicle registration certificates, which must undergo optical character recognition (OCR) recognition and anonymization processing by the platform), educational certificate information, and part-time income records (such as service settlement records for freelancers). This self-uploaded data can serve as supplementary credit dimensions, further improving the user's data profile and enhancing the comprehensiveness and accuracy of credit assessment.

[0041] The types of pending data sent from different data sources can vary. For example, pending data provided by content-based social platforms may include: user account registration information (after anonymization), content consumption preferences, live streaming reward records, and virtual asset transactions and fulfillment status within the platform (such as e-commerce order payment and return records). Pending data provided by third-party payment platforms may include: basic payment account information (after anonymization), transfer records, bill fulfillment status, quick payment binding and usage frequency, etc. Pending data provided by real estate transaction platforms may include: user real estate transaction intention records, housing rental / sale contract fulfillment status, rent / housing payment records and overdue information, etc. Pending data provided by travel service platforms may include: user travel order completion status, fare payment records, and changes in platform credit scores (related to fulfillment behavior), etc.

[0042] It should be noted that sensitive personal information contained in the data to be processed can be encrypted or anonymized beforehand. Without exposing the original plaintext, the data matching task can be completed collaboratively by multiple data sources based on Multi-Party Secure Computation (MSC). During this process, each participating party can only obtain the authorized final results or aggregated metrics, and cannot access the original input data of other parties.

[0043] The aforementioned entity objects can be set according to actual needs, including but not limited to: Different entity objects correspond to different users. At this time, the matching result set corresponding to each entity object contains various types of data of that user, which are used to represent the user's data profile. For example, a user's personal credit profile is built through personal credit data to provide a basis for application scenarios such as credit approval and service rights assessment.

[0044] For example, by verifying users' occupational information, income-related consumption patterns, and asset allocation data (such as financial purchase records), a user's economic capability profile can be constructed, thereby assisting insurance institutions in setting insurance premium rates, credit institutions in optimizing credit limit approvals, and high-end service platforms in segmenting customers.

[0045] Different entities correspond to different hotels. At this time, the matching results set for each entity includes the hotel's operating license information, customer booking fulfillment records, online review data, payment and settlement behavior, and cooperation credit history with the platform. This is used to characterize the hotel's service credit and operational stability profile. For example, by analyzing the hotel's order cancellation rate, customer complaint handling timeliness, and compensation fulfillment status on multiple booking platforms, a service fulfillment credit profile can be constructed to provide decision support for online travel platforms (OTAs) to dynamically adjust recommendation weights and verify deposit-free check-in qualifications.

[0046] For example, by integrating the hotel's payment settlement timeliness data (such as the monthly settlement completion rate with OTA platforms and whether there are overdue payment records), the validity period of the operating license and the annual inspection compliance status, and the hygiene and safety related scores in online reviews (such as the hygiene-related positive review rate in the past 3 months), a profile of its operational health can be constructed, thereby providing a basis for OTA platforms to adjust settlement periods.

[0047] Different entity objects correspond to different regions. At this time, the matching result set corresponding to each entity object contains various types of data of the region, which are used to characterize the data profile of the region. For example, by constructing the regional consumption level, consumption structure and other characteristics through regional consumption data, a consumption data profile of the region can be built, thereby providing a reference for regional economic analysis, business layout planning and so on.

[0048] For example, by using data on the distribution of public service resources within the region (such as the number of schools and hospitals), the residents' credit compliance rate, and data on the performance of infrastructure construction, a profile of the region's people's livelihood can be constructed, thereby supporting the government's planning of livelihood projects, evaluation of the implementation of public welfare projects, and optimization of social governance efficiency.

[0049] Different entity objects correspond to different enterprises. At this time, the matching results for each entity object contain the enterprise's operation and credit-related data, which are used to represent the enterprise's data profile. For example, a business credit profile of an enterprise can be built through tax records, loan performance, supply chain transaction data, etc., thereby supporting financial institutions to carry out business such as corporate credit granting, supply chain finance, or risk monitoring.

[0050] For example, by using data such as corporate R&D investment, intellectual property registration records, and compliance innovation qualification certifications, a profile of a company's innovative development can be constructed, thereby providing a reference for investment institutions' project evaluation and market partners' technology cooperation decisions.

[0051] S202: Perform at least two rounds of data matching on each of the data to be processed to obtain each matching result set for use by downstream business systems; wherein, each matching result set is used to represent a data profile of an entity object.

[0052] Furthermore, the formats of the data to be processed sent from different data sources often differ. For example, repayment records sent by bank loan systems may record dates in "YYYYMMDD" format, while consumption records from e-commerce platforms may use "MM / DD / YYYY" format. In public utility payment records, the "ID number" field may be named "Document Number," and after anonymization, only the first 6 and last 4 digits are retained (e.g., "110101xxxxxxxx1234"). In contrast, user information provided by telecommunications operators may name the "ID number" "Identity Identifier," with the anonymization rule retaining only the first 8 and last 4 digits (e.g., "1101011912xxxx1234"). The credit reporting platform's server can receive the data to be processed from various data sources and preprocess it to convert it into a standardized format.

[0053] The credit reporting platform's server can employ various methods to preprocess the data to be processed, and this manual does not impose any restrictions on these methods. Examples include: field normalization (e.g., mapping "document number" and "identity identifier" to the "ID card number" field), format standardization (e.g., converting all dates to "YYYY-MM-DD" format and amounts to "yuan" unit with two decimal places), and anonymization rule adaptation (e.g., extracting fixed-digit features for matching based on different anonymized ID card numbers, such as the first 6 digits of the administrative region code and the last 4 digits of the sequence code). These operations aim to convert the data to be processed into a standardized format, providing a unified data foundation for subsequent multiple rounds of data matching and reducing matching errors caused by format differences.

[0054] Furthermore, in practical applications, the types of data to be processed sent by different data sources are often different. For example, data source A sends structured document PDFs, data source B sends semi-structured data (such as JSON data), data source C sends image data, and data source D sends voice data, etc.

[0055] At this point, the credit reporting platform's server can also select a model from a preset model pool that matches the type of data to be processed for each type of data, and use the target model to convert the data to be processed into a standardized format.

[0056] For example, if the data to be processed is structured document data, a pre-trained Large Language Model (LLM) can be selected from the model pool as the target model. The target model can then be guided by prompts (such as "extract the fields 'applicant's name', 'ID number', and 'loan amount' from the document, output them in JSON format, and convert the dates to YYYY-MM-DD") to quickly parse the structured fields and output the data to be processed in a standardized format.

[0057] For example, if the data to be processed is semi-structured data, a lightweight structure parsing model based on Transformer can be selected as the target model. This allows the target model to automatically learn the field-level features of different data sources, while unifying field naming through semantic similarity algorithms (e.g., mapping "payment date" and "transaction date" to "transaction time") and outputting data to be processed in a standardized format.

[0058] For example, if the data to be processed is unstructured, a conditional random field fusion model can be selected as the target model to extract text information (including handwriting recognition) from the image. Then, the recognition error is corrected based on contextual semantics (e.g., "2023.13.05" is corrected to "2023-12-05"). At the same time, key fields such as "payer" and "payment amount" are located through entity recognition algorithms, and finally, the data to be processed in a standardized format is output.

[0059] In addition, the aforementioned preprocessing operations may also include quality checks and compression of the data to be processed. Quality checks may include verification of data integrity (e.g., whether key fields are missing), consistency (e.g., whether ID numbers and names logically match), timeliness (e.g., whether the recording time is within a reasonable range), and compliance (e.g., whether it complies with anonymization rules or privacy protection requirements). For data that does not meet quality requirements, it can be marked, removed, or a completion / correction process can be triggered (i.e., sending a data completion / correction request to the data source so that the data source responds to the request and re-uploads the completed / corrected data).

[0060] Meanwhile, to improve storage performance, the credit reporting platform's server can also compress the data to be processed. For example, it can use general compression algorithms (such as GZIP, Snappy, etc.) to perform lossless compression on structured or semi-structured data, or simplify the encoding of redundant fields while ensuring no information loss. These preprocessing operations help improve the accuracy of subsequent data matching and the overall system operating efficiency, while reducing resource consumption.

[0061] Furthermore, after standardizing each piece of data to be processed, the server of the credit reporting platform can perform at least two rounds of data matching to obtain each matching result set for downstream business systems to call. For ease of understanding, the process of any round of data matching is explained in detail below, including steps S2021 to S2022.

[0062] S2021: Based on the preset credibility level corresponding to each data source, determine the target data required for this round of data matching from the data to be processed. The higher the round of data matching, the higher the credibility level of the data source corresponding to the determined target data.

[0063] Different confidence levels can be preset for different data sources. Then, in each round of data matching, the target data required for that round of data matching can be determined from each data to be processed based on the confidence level corresponding to the different data sources.

[0064] Specifically, for any given round of data matching, the higher the round number, the higher the credibility level of the data source corresponding to the identified target data. For example, if this is the first round of data matching, then the data to be processed sent by the data sources with the lowest credibility levels can be used as the target data for this round. Subsequently, for each additional round of data matching, the credibility level of the corresponding target data increases by a specified value (e.g., for each additional round of data matching, the credibility level of the corresponding target data increases by one level).

[0065] In the above, the credibility level corresponding to different data sources can be determined in advance based on the accuracy of each piece of data to be processed that has been sent historically from different data sources.

[0066] Of course, for the first time sending data to be processed from various data sources, the initial credibility level can be determined based on the data source's qualification and compliance, the completeness of data fields, and its conformity with standards for different application scenarios. For example, if a new data source has credit reporting business qualifications recognized by a designated certification authority, and the completeness of the data it sends is high, and the format of the data fully complies with the standardization requirements of the credit reporting platform, then its initial credibility level can be determined as high. If the new data source lacks clear compliance qualifications, or if there are many missing data fields or the format seriously deviates from the standard, then the initial credibility level is determined as low, and the level will be dynamically adjusted after subsequent data undergoes multiple rounds of matching and verification of accuracy.

[0067] S2022: Take each target data and each matched data contained in the matching result set output by the previous round of data matching as each input data, and input each input data into the preset matching model, so that the matching model can identify input data belonging to the same entity object from each input data and merge them into the same matching result set, so as to obtain the matching result set output by this round of data matching.

[0068] Figure 3 This is a schematic diagram of any round of data matching provided in an exemplary embodiment.

[0069] Combination Figure 3 As can be seen, after determining the target data required for this round of data matching, the credit reporting platform's server can select the matching model needed for this round of data matching from a pre-set model pool. Then, it can input the target data and the matched data from the previous round's matching result set as input data into the matching model. The matching model then identifies input data belonging to the same entity from the input data (including the target data required for this round of data matching and the matched data from the previous round's matching result set) and merges them into the same matching result set, thus obtaining the matching result set output for this round of data matching. Here, merging can refer to the logical aggregation of multiple input data identified as belonging to the same entity to form a unified, structured data set.

[0070] To facilitate understanding, the following will use the aforementioned entity object, the user, as an example to explain the matching process in detail.

[0071] Specifically, since all data uploaded by each data source has undergone pre-sensitization processing, to ensure user privacy and security and avoid the risk of identity data leakage, for at least some of the data sources, before sending the data to be processed to the credit reporting platform's server, the user's identity data contained in at least some of the data to be processed can be converted into desensitized virtual identity identifiers according to a specified encryption algorithm. Here, the specified encryption algorithm can be an encryption algorithm jointly determined and used by the credit reporting platform and each data source in advance.

[0072] For example, before sending a user's credit record, a data source uses a hash algorithm (such as SHA-256) to encrypt the user's real ID number "11010119900101**34" to generate a fixed-length string "a8b2d3e5f7... (hash characters)" as a virtual identity identifier. Since all data sources use the same encryption rules (such as the same hash salt value and encryption key), the real identity identifier of the same user in different data sources is converted into the same virtual identity identifier. This not only protects privacy but also enables the matching of the same user's data across data sources on the credit reporting platform based on the virtual identity identifier.

[0073] Based on this, the server of the credit reporting platform can input each input data into a preset matching model. The matching model can determine the entity object whose virtual identity is consistent with the virtual identity contained in the input data for each input data containing a virtual identity. The entity object that matches the input data is then assigned to the matching result set corresponding to the entity object.

[0074] In addition, since some data sources may not have used the specified encryption algorithm to generate the corresponding virtual identity, the matching model can also perform semantic analysis on each input data that does not contain the virtual identity to determine the semantic correlation between the input data and the data already included in the matching result set corresponding to each entity object. If the semantic correlation is greater than a preset threshold, the input data is assigned to the matching result set corresponding to that entity object. For example, if the entity object (user) resides in city A, it can be determined that the semantic correlation between the logistics record data with a delivery address in city B and the entity object is significantly lower than the preset threshold, thus determining that the two are not related. For example, if the matching result set of an entity object (user) has recorded its frequently used contact number as "138XXXX5678", and the emergency contact name associated with this number in the historical data is "Zhang X", then when an input data without a virtual identity contains "emergency contact: Zhang X, contact number: 138XXXX5678", the matching model can determine through semantic analysis that the semantic correlation between the two is significantly higher than the preset threshold, and thus classify the input data into the matching result set corresponding to the user.

[0075] It should be noted that the matching model described above for matching each input data containing the virtual identity identifier and the matching model described above for matching each input data not containing the virtual identity identifier can be the same model or different models. For example, a rule-based matching model based on precise identifier alignment can be used to match each input data containing the virtual identity identifier, while a large language model can be used to match each input data not containing the virtual identity identifier.

[0076] Furthermore, in practical applications, there may be situations where multiple data sources for the same entity object provide inconsistent values ​​for a certain key attribute (such as education level, occupation information, monthly income level, etc.). Based on this, for each entity object's matching result set, if the matching model determines that there are two different input data for any specified attribute of the entity object in the matching result set, then the input data with the highest confidence level of the corresponding data source is selected from the input data corresponding to the specified attribute and used as the standard input data corresponding to the specified attribute.

[0077] The specified attributes mentioned above are used to characterize the features or state of an entity object. Examples include: education level, occupational information, and monthly income level.

[0078] To facilitate understanding, the following example illustrates how a matching model determines the standard input data corresponding to a specified attribute: If the aforementioned entity is user A, and the aforementioned specified attribute is user A's monthly income, then if each input data contains pending data sent by data source A to represent user A's monthly income as 50,000 (unit: yuan / year), and each input data also contains pending data sent by data source B to represent user A's monthly income as 10,000, and pending data sent by data source C to represent user A's monthly income as 500.

[0079] Assumptions: Data source A: credibility level 5, Data source B: credibility level 8, Data source C: credibility level 5.

[0080] At this point, the matching model can select the monthly income data provided by data source B with the highest credibility level as the standard input data for the specified attribute.

[0081] In addition, the matching model can determine the credibility weight of each input data for the specified attribute based on the credibility level of each data source corresponding to the input data and the weight coefficients corresponding to different credibility levels. Then, it can select the input data with the highest credibility weight from all the input data as the standard input data for the specified attribute.

[0082] For example: If the above entity object is user A, and the above specified attribute is user A's monthly income, then if each input data contains pending data sent by data source A to represent user A's monthly income as 50,000, and each input data also contains pending data sent by data source B to represent user A's monthly income as 10,000, pending data sent by data source C to represent user A's monthly income as 50,000, and pending data sent by data source D to represent user A's monthly income as 50,000.

[0083] Assumptions: Data source A: credibility level 5, Data source B: credibility level 8, Data source C: credibility level 5, Data source D: credibility level 4.

[0084] At this point, the specified attribute contains two types of input data: input data with a value of 50000 and input data with a value of 10000. The input data with a value of 50000 corresponds to three data sources (data source A, data source C, and data source D), while the input data with a value of 10000 corresponds to only one data source (data source B).

[0085] Based on this, the confidence weights for the input data with a value of 50000 determined by the matching model are: (5+5)*0.6+4*0.5=8 for data source A, and 8*0.9=7.2 for input data with a value of 10000. Here, 0.6 is the weight coefficient corresponding to confidence level 5, 0.5 is the weight coefficient corresponding to confidence level 4, and 0.9 is the weight coefficient corresponding to confidence level 8.

[0086] Therefore, although data source B has a higher individual credibility level than any of data sources A, C, or D, its overall credibility weight of 8 is higher than the credibility weight of 7.2 corresponding to the input data with a value of 10000 because the input data with a value of 50000 is supported by multiple data sources. Therefore, the matching model will select the input data with a monthly income of 50000 as the standard input data for this specified attribute (user A's monthly income). The weight coefficients for each credibility level mentioned above can be set according to actual needs; the weight coefficients above are for illustrative purposes only.

[0087] Of course, in practical applications, the matching model can generate a label for each input data based on the credibility level of the data source corresponding to each input data, or based on the credibility weight corresponding to each input data, so as to label each input data.

[0088] For example, the matching model can generate labels to mark the input data for the specified attribute provided by the data source with the highest corresponding confidence level as "high confidence", while marking the input data for the specified attribute provided by other data sources as "to be verified" or "low confidence", etc.

[0089] Similarly, the matching model used here to determine any specified attribute of the standard input data can be a separately set matching model dedicated to consistency detection, or it can be the same model as the aforementioned matching model used to merge input data belonging to the same entity object into the same matching result set.

[0090] Furthermore, since the original input data is usually characterized by high dimensionality, sparseness, and heterogeneity, directly using it for business decision-making may result in information redundancy or semantic ambiguity. In this specification, the server of the credit scoring platform can also input each input data into a preset matching model, so that the matching model can identify input data belonging to the same entity object from each input data and merge them into the same matching result set to obtain the basic matching result set of each entity object. For each entity object, based on the basic matching result set corresponding to the entity object, the profile feature data of the entity object is generated. Based on the profile feature data corresponding to the entity object and the basic matching result set, the matching result set output after this round of data matching is obtained.

[0091] The aforementioned portrait feature data are derived features of the entity object obtained by feature extraction or feature prediction from the basic matching result set.

[0092] To facilitate understanding, the following will use the aforementioned entity object, the user, as an example to explain the above-mentioned profile feature data in detail: When the aforementioned entity is a user, the matching model performs same-person identification on each input data to merge the input data belonging to the same user into the matching result set corresponding to that user, thereby obtaining the basic matching result set corresponding to that user. Subsequently, the matching model can also perform regional identification, income prediction, preference prediction, number of children prediction, education model construction, hotel model construction, airport model construction, etc., based on the basic matching result set corresponding to that user, to generate profile feature data corresponding to that user to supplement the basic matching result set, thereby obtaining the final matching result set output by this round of data matching.

[0093] For example, by inputting data such as the user's permanent IP address, communication base station location, social security payment location, and utility bill payment address, the region can be identified to determine the user's profile features such as "city of residence" and "level of residential stability". By inputting data such as salary deposit records, consumption amount distribution, credit card limit utilization rate, and financial asset change trends, income range prediction is performed to determine user profile characteristics such as "estimated monthly income" and "income credibility level". By using information such as academic qualification verification, professional qualification certificate application records, and online learning platform behavior, a user's educational background and lifelong learning ability model can be constructed to determine the user's profile characteristics such as "confidence level of education level" and "distribution of skills fields". By analyzing the star rating, price range, location, and frequency of hotel bookings at different times, a user hotel preference model is constructed to determine user profile characteristics such as "business travel tendency" and "high-end accommodation preference index". By analyzing flight purchase records, airport security check frequency, airline membership levels, and destination distribution, an airport usage and travel model is constructed to identify user profile characteristics such as "high-frequency flyers" and "international travel activity." By using education payment records, children's medical treatment information, and family package communication data, family structure can be predicted to determine profile characteristics such as "whether or not they have children" and "the age range of their children".

[0094] Furthermore, for each round of data matching, after the matching model obtains the matching result set corresponding to each entity object, it can also determine a corresponding label for each input data contained in the matching result set corresponding to that entity object, in order to identify the credibility status of the input data (e.g., high credibility, pending verification, low credibility, etc.). In addition, the labels determined by the matching model for each input data can also be used to characterize information such as the source attribute, timeliness, completeness, or business relevance of the input data. For example, labels such as "high credibility," "pending verification," "historical archive," "core identity type," "behavioral trajectory type," and "conflict to be resolved" can be assigned to the input data.

[0095] In practical applications, the matching result set used in each round of data matching can be either the entire matching result set output from the previous round of data matching, or a matching result set composed of at least a portion of the input data whose labels are specified (e.g., "highly reliable" or "confirmed").

[0096] It should be noted that the matching result set used in the first round of data matching can be an empty set or a set of historical matching results determined by the credit reporting platform's server during the historical data matching process.

[0097] Furthermore, since different downstream business systems may have different requirements for data accuracy (for example, a credit approval system needs a highly reliable matching result set obtained after multiple rounds of data matching, while a basic credit query only needs to use the matching result set obtained after one round of data matching), this specification retains the matching result sets obtained after each round of data matching. This is so that upon receiving a data request from a downstream business system, the system can select and return the matching result sets output after the corresponding rounds of data matching based on the filtering conditions carried in the data request. Specifically, as follows... Figure 4 As shown.

[0098] Figure 4 This is a schematic diagram of multi-round data matching provided in an exemplary embodiment.

[0099] Combination Figure 4 It can be seen that the matching result sets corresponding to each entity object obtained after each round of data matching can serve as input data for the next round of data matching, thereby supporting entity merging with higher accuracy and reliability. For example, in the second round of data matching, the matching result set output by the first round of data matching is introduced and merged with the data from a new round of high-reliability data sources, thus gradually converging to more accurate entity boundaries and attribute values.

[0100] On the other hand, it can serve as the output storage for this round of data matching, in order to respond to the differentiated needs of different downstream business systems for data quality and timeliness.

[0101] It's important to note that different data types, data sources, and application scenarios place significantly different demands on matching models. For example, different data types (such as text, numbers, and dates) may require different matching algorithms or models to improve matching accuracy. Different types of data sources have different structural characteristics and noise levels, thus necessitating the use of different matching algorithms or models. Furthermore, different application scenarios (such as personal credit assessment, corporate risk control, and market trend analysis) have varying requirements for the matching result set (e.g., in personal credit assessment, the focus is on capturing and identifying subtle differences between individuals; while in corporate risk control, the focus may be more on the complex relationship networks between related companies), thus requiring different matching algorithms or models.

[0102] Based on this, after determining the target data required for any round of data matching, the server of the credit reporting platform can also select at least one matching model from the matching models contained in the preset model pool as the target matching model according to the matching conditions of that round of data matching.

[0103] The matching conditions mentioned above include at least one of the following: the data type of the target data, the data source type to which the target data belongs, and the preset application scenario type.

[0104] Of course, since a single matching model may have blind spots or insufficient generalization ability when facing highly heterogeneous, complex noise or cross-domain data, the credit reporting platform's server can also select a target model group from the matching models included in the model pool, input each input data into each target matching model included in the target matching model group, and for each target matching model, identify input data belonging to the same entity object from each input data and merge them into the same matching result set.

[0105] For each target model group, there are at least two matching models.

[0106] Based on this, the credit reporting platform's server can select a target model group from the matching models in the preset model pool that meet the preset matching conditions, based on different preset matching conditions.

[0107] The matching models included in the aforementioned model pool can include: Large Language Model (LLM), Random Forest model, rule-based deterministic matching model, and deep learning-based semantic matching model, etc. Within this model pool, for the same type of model, multiple models can be provided with different parameter sizes, training data sources, feature input dimensions, inference accuracy, and latency requirements.

[0108] For example, for large language models (LLM), the above model pool can simultaneously deploy lightweight distilled versions (such as TinyBERT and DistilLLaMA) for high-concurrency, low-latency scenarios, as well as full-parameter versions (such as LLaMA-2-13B and ChatGLM3) for high-precision, complex semantic alignment tasks.

[0109] Of course, in real-world applications, the types and quantities of matching models included in the model pool can be added or deleted according to actual needs.

[0110] In addition, the server of the business platform can also fuse the matching result sets output by each target matching model according to the weight coefficients of each target matching model in the target matching model group under the matching conditions of this round of data matching, so as to obtain the matching result set output by this round of data matching.

[0111] The weight coefficients of any matching model under different matching conditions are determined based on the accuracy of the historical matching result set output by the matching model under those matching conditions.

[0112] In the above content, the business platform's server can use various methods to fuse the matching result sets output by each target matching model based on the weight coefficients of each target matching model in the target matching model group under the matching conditions of this round of data matching.

[0113] For example, to determine whether any two input data belong to the same entity object, the business platform's server can perform weighted voting on the matching result set output by each target matching model based on the weight coefficients of each target matching model in the target matching model group under the matching conditions of this round of data matching. If the weighted score of "belonging to the same entity object" exceeds a preset threshold, then it can be determined that any two input data can be merged into the same matching result set.

[0114] For example, from the target matching model group, the target matching model with the highest weight coefficient is selected as the main matching model. The matching result set output by the main matching model can be used as the main matching result set. Input data with labels such as "high confidence" or "confirmed" in the main matching result set can be retained. Input data with labels such as "low confidence" or "pending confirmation" in the main matching result set can be cross-validated and supplemented based on the matching result sets output by other matching models. For example, if a low confidence input data is merged into the same entity in the matching result sets output by multiple other matching models, and its attribute conflict is small, its corresponding label can be changed to "high confidence" and included in the final result. If there are significant discrepancies between different other matching models, its corresponding label can be changed to "requires manual review" to further verify its attribution.

[0115] Therefore, in this specification, each matching result set includes the matching result sets output in different rounds of data matching. For each round of data matching, each matching result set output after that round of data matching includes the matching result sets output by different matching models or combinations of different matching models contained in the preset model pool.

[0116] In summary, this specification demonstrates that the credit reporting platform's server-side mechanism can, on the one hand, improve the credibility level of the data source corresponding to the target data in each round, ensuring that high-credibility data plays a crucial calibration role in subsequent matching and reducing the risk of mismatches caused by inconsistent formats and differences in identification rules of low-credibility data. On the other hand, each round of matching is based on the results of the previous round and incorporates new data with higher credibility, forming an iteratively optimized matching chain. This effectively links scattered data of the same subject, reducing the problem of missed matches caused by non-standard data identification. Ultimately, the entity object data profile formed through multiple rounds of matching significantly improves the completeness and accuracy of personal data on the credit reporting platform.

[0117] To facilitate understanding, the following provides a detailed explanation of the business execution methods performed by the server-side of the aforementioned credit reporting platform during the business execution process, as detailed below. Figure 5 As shown.

[0118] Figure 5 This is a flowchart illustrating a business execution method provided in an exemplary embodiment, including the following steps: S500: Receive a data retrieval request; the data retrieval request includes filtering conditions, which include at least one of the following: specifying a credibility level, specifying a matching model type, and specifying a business scenario.

[0119] In actual business operations, when a data caller (e.g., any user, housing rental platform, bank credit system, e-commerce platform, etc.) needs to access credit data to support its business operations, it can first initiate a data call request to the credit reporting platform's server through its client or business platform server. This request not only includes the data caller's identity identifier (such as user ID, organization code, etc.) but can also carry filtering conditions to specify the quality and scope of the desired matching result set.

[0120] The above-mentioned screening conditions may include, but are not limited to, at least one of the following: specifying a credibility level (e.g., specifying high credibility, specifying medium credibility, specifying basic credibility, etc.), specifying a matching model type (e.g., specifying a matching result set determined by a large language model, specifying a matching result set determined by a random forest model, etc.), and specifying a business scenario (e.g., "credit approval", "deposit-free rental", "insurance underwriting", etc.).

[0121] S502: Determine the target matching result set that matches the filtering conditions from the pre-processed matching result sets and return it.

[0122] Upon receiving the data request, the credit reporting platform's server can parse the identifier of the data caller carried within it and, based on this identifier, determine the target trusted domain corresponding to the data call method from multiple pre-defined trusted domains. These trusted domains are logically isolated data storage and access areas pre-created by the credit reporting platform for different data callers or groups of data callers, ensuring that each data caller can only access authorized data bound to their identity, thus meeting data security and privacy compliance requirements.

[0123] The aforementioned data caller group can refer to a collection of data callers that have the same business attributes, security level, management requirements, or data usage permissions.

[0124] For example, multiple platforms in the same industry (such as multiple commercial banks, insurance companies, or consumer finance companies) can be grouped into the same group because they are highly consistent in their use of personal credit data, the scope of fields, and compliance standards.

[0125] For example, multiple subsidiaries belonging to the same group or parent company, although independent legal entities, can be grouped into the same group if they are managed uniformly in terms of data management and access control policies.

[0126] Subsequently, the credit reporting platform's server retrieves and matches data from its pre-defined, multi-round data matching sets, based on the filtering criteria in the data retrieval request. For example, if the request specifies "credibility level = 3" and "business scenario = credit approval," the credit reporting platform's server will select the matching result set that has undergone three or more rounds of matching as the target matching result set and return it.

[0127] Furthermore, after determining the target matching result set, the credit reporting platform's server can store or map it into the aforementioned target trusted space domain, and return a successful call response and data access credentials to the requester. Downstream business systems can then securely and efficiently read the required credit data within this target trusted space domain for credit assessment, risk control, or service access and other business operations.

[0128] In practical applications, the credit reporting platform's server can also, after responding to each received data retrieval request and returning the corresponding target matching result set to the data caller, generate a data retrieval record based on the data caller's identifier, data retrieval time, the unique identifier of the target matching result set, and the user's authorization credential number. The data retrieval record is then sent to each data provider of the data contained in the aforementioned target matching result set, thereby making data usage behavior transparent and traceable.

[0129] As can be seen from the above, the credit reporting platform's server can meet the differentiated data quality needs of different businesses while ensuring the privacy of user data through an on-demand supply, tiered authorization, spatial isolation, and secure and controllable credit data service model.

[0130] Figure 6 This is a schematic structural diagram of a device provided in an exemplary embodiment. For example... Figure 6As shown, device 600 mainly consists of a communication interface 602, a user interface 604, a processor 606, and a data storage 608. These components are interconnected and communicate with each other via a system bus, network, or other connection mechanism 610. The communication interface 602 enables device 600 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface 602 may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface 602 can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface 602 can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface 602 may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces.

[0131] User interface 604 includes receiving user input and providing output to the user. Therefore, user interface 604 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. User interface 604 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, user interface 604 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, device 600 may support remote access from other devices via communication interface 602 or another physical interface (not shown). User interface 604 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. User interface 604 may also be configured as a display device for rendering or displaying text fragments.

[0132] Processor 606 may contain one or more general-purpose processors and / or special-purpose processors.

[0133] Data storage 608 may include one or more volatile and / or non-volatile storage components and may be integrated wholly or partially with processor 606. Data storage 608 may include removable and non-removable components.

[0134] Processor 606 is capable of executing program instructions 618 (e.g., compiled or uncompiled program logic and / or machine code) stored in data storage 608 to perform the various functions described herein. Data storage 608 may contain a non-transitory computer-readable medium on which program instructions are stored, which, when executed by device 600, enable device 600 to perform any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Execution of program instructions 618 by processor 606 may result in processor 606 using data 612.

[0135] For example, program instructions 618 may include an operating system 622 (e.g., an operating system kernel, device drivers, and / or other modules) installed on device 600 and one or more applications 620 (e.g., a browser, social application, or game application). Similarly, data 612 may include operating system data 616 and application data 614. Operating system data 616 is primarily accessible to the operating system 622, while application data 614 is primarily accessible to one or more applications 620. Application data 614 may reside in a file system visible or hidden from the user of device 600.

[0136] Application 620 can communicate with operating system 622 through one or more application programming interfaces (APIs). These APIs help application 620 read and / or write application data 614, transmit or receive information via communication interface 602, receive or display information on user interface 604, etc.

[0137] In some terminology, application 620 may be simply referred to as "app". Furthermore, application 620 can be downloaded to device 600 through one or more online app stores or app markets. However, applications can also be installed on device 600 in other ways, such as through a web browser or a physical interface on device 600 (e.g., a USB port).

[0138] Please refer to Figure 7 A data matching device can be applied to, for example Figure 6 The device shown is used to implement the technical solution described in this specification.

[0139] The receiving module 701 is used to receive the data to be processed sent by each data source; The matching module 702 is used to perform at least two rounds of data matching on each piece of data to be processed to obtain each matching result set for use by downstream business systems. Each matching result set is used to represent a data profile of an entity object. In any round of data matching: based on the preset credibility level corresponding to each data source, the target data required for this round of data matching is determined from each piece of data to be processed. The higher the round of data matching, the higher the credibility level of the data source corresponding to the determined target data. The target data and the matched data contained in the matching result set output by the previous round of data matching are used as input data. The input data are input into a preset matching model so that the matching model can identify input data belonging to the same entity object from each input data and merge them into the same matching result set to obtain the matching result set output by this round of data matching.

[0140] Optionally, each piece of data to be processed has been de-identified in advance, and at least a portion of the data to be processed contains the virtual identity identifier of the corresponding user.

[0141] Optionally, the matching module 702 is specifically used to input each input data into a preset matching model, so that the matching model, for each input data containing the virtual identity identifier, determines the entity object that matches the input data based on the virtual identity identifier contained in the input data, and assigns the input data to the matching result set corresponding to the entity object.

[0142] Optionally, the matching module 702 is specifically used to input each input data into a preset matching model, so that the matching model performs semantic analysis on each input data that does not contain the virtual identity identifier, to determine the semantic correlation between the input data and the data already contained in the matching result set corresponding to each entity object, and when the semantic correlation is determined to be greater than a preset threshold, the input data is assigned to the matching result set corresponding to the entity object.

[0143] Optionally, for each entity object's matching result set, if the matching model determines that there are two different input data corresponding to any specified attribute of the entity object in the matching result set, then the input data with the highest confidence level of the corresponding data source is selected from the input data corresponding to the specified attribute and used as the standard input data corresponding to the specified attribute; the specified attribute is used to characterize the features or state of the entity object.

[0144] Optionally, the matching module 702 is specifically used to input the input data into a preset matching model, so that the matching model identifies input data belonging to the same entity object from the input data and merges them into the same matching result set to obtain a basic matching result set for each entity object. For each entity object, based on the basic matching result set corresponding to the entity object, it generates portrait feature data of the entity object. Based on the portrait feature data corresponding to the entity object and the basic matching result set, it obtains the matching result set output after this round of data matching. The portrait feature data is the derived feature of the entity object obtained by feature extraction or feature prediction of the basic matching result set.

[0145] Optionally, the matching module 702 is specifically used to select at least one matching model from the matching models included in the preset model pool as the target matching model according to the matching conditions of the data matching in this round; the matching conditions include at least one of the following: the data type of the target data, the data source type to which the target data belongs, and the preset application scenario type; and input the input data into the target matching model so that the target matching model can identify input data belonging to the same entity object from each input data and merge them into the same matching result set.

[0146] Optionally, the matching module 702 is specifically used to select a target model group from the matching models included in the model pool, wherein the target model group includes at least two matching models; input the input data into each target matching model included in the target matching model group; and for each target matching model, identify input data belonging to the same entity object from the input data and merge them into the same matching result set through the target matching model.

[0147] Optionally, the matching module 702 is further configured to fuse the matching result sets output by each target matching model according to the weight coefficients of each target matching model included in the target matching model group under the matching conditions of this round of data matching, so as to obtain the matching result set output by this round of data matching; wherein, the weight coefficient of any matching model under different matching conditions is determined according to the accuracy of the historical matching result set output by the matching model under the matching conditions.

[0148] Optionally, each matching result set includes the matching result sets output in different rounds of data matching; for each round of data matching, each matching result set output after that round of data matching includes the matching result sets output by different matching models or combinations of different matching models contained in a preset model pool.

[0149] Please refer to Figure 8 A business execution device can be applied to, for example Figure 6The device shown is used to implement the technical solution described in this specification.

[0150] The request receiving module 801 is used to receive a data call request; the data call request includes filtering conditions, which include at least one of the following: specifying a credibility level, specifying a matching model type, and specifying a business scenario. The filtering module 802 is used to determine and return the target matching result set that matches the filtering conditions from the pre-processed matching result sets.

[0151] Optionally, the target matching result set is sent to a target trusted space domain for storage, so that downstream business systems can call it in the target trusted space domain; the target trusted space domain is determined from preset trusted space domains based on the identifier of the data caller contained in the data call request; the trusted space domain is a logically isolated data storage and access space created for different data callers or groups of data callers.

[0152] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0153] Based on the same concept as the methods described above, this specification also provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor performs the steps of the method as described in any of the above embodiments by executing the executable instructions.

[0154] Based on the same concept as the methods described above, this specification also provides a computer-readable storage medium having computer instructions stored thereon that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0155] Based on the same concept as the methods described above, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the methods as described in any of the above embodiments.

[0156] What those skilled in the art will understand is: In this specification, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitation, the presence of additional identical or equivalent elements in a process, method, product, or apparatus that includes said elements is not excluded.

[0157] In this specification, “a,” “an,” and “the” do not specifically refer to the singular, but may also include the plural.

[0158] In this specification, ordinal numbers such as "first," "second," etc., do not necessarily indicate order; they are often used to distinguish between objects. For example, "first server" and "second server" usually refer to two servers. To differentiate between these two servers, they are described as "first server" and "second server." Of course, sometimes these two servers may be the same server.

[0159] In this specification, unless explicitly stated otherwise, "receiving and sending data" does not necessarily mean direct receiving and sending; it can also mean indirect receiving and sending. For example, A receiving data sent by B can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, B sending data to A can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.

[0160] In this specification, unless explicitly stated otherwise, the relationships between structures can be direct or indirect. For example, when describing "A is connected to B," unless it is explicitly stated that A and B are directly connected, it should be understood that A can be directly connected to B or indirectly connected to B. Similarly, when describing "A is on top of B," unless it is explicitly stated that A is directly above B (AB is adjacent and A is above B), it should be understood that A can be directly above B or indirectly above B (AB is separated by other elements, and A is above B). And so on.

[0161] This specification uses specific terms to describe embodiments thereof. Terms such as "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described herein, as well as the features of those different embodiments or examples, without contradiction.

[0162] Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is only one of many possible execution orders and does not represent the only execution order. Therefore, when the claims involve method steps, any changes or adjustments to the order of such steps, or the parallelism between steps, are also within the scope of protection of the claims.

Claims

1. A data matching method, the method comprising: Receive data to be processed from various data sources; At least two rounds of data matching are performed on each piece of data to be processed to obtain each matching result set for use by downstream business systems; wherein, each matching result set is used to represent a data profile of an entity object; In any round of data matching: Based on the preset credibility level of each data source, the target data required for this round of data matching is determined from each data to be processed. The higher the round of data matching, the higher the credibility level of the data source corresponding to the determined target data. Each target data and each matched data included in the matching result set output from the previous round of data matching are used as input data. Each input data is then input into a preset matching model so that the matching model can identify input data belonging to the same entity object from each input data and merge them into the same matching result set to obtain the matching result set output from this round of data matching.

2. The method as described in claim 1, wherein each piece of data to be processed has been de-identified in advance, and at least a portion of the data to be processed contains the virtual identity identifier of the user corresponding to it.

3. The method as described in claim 2, wherein the input data is input into a preset matching model, so that the matching model identifies input data belonging to the same entity object from the input data and merges them into the same matching result set, specifically including: Each input data is input into a preset matching model, so that the matching model, for each input data containing the virtual identity identifier, determines the entity object that matches the input data based on the virtual identity identifier contained in the input data, and assigns the input data to the matching result set corresponding to the entity object.

4. The method as described in claim 2, wherein the input data is input into a preset matching model, so that the matching model identifies input data belonging to the same entity object from the input data and merges them into the same matching result set, specifically including: Each input data is input into a preset matching model, so that the matching model performs semantic analysis on each input data that does not contain the virtual identity identifier, determines the semantic correlation between the input data and the data already contained in the matching result set corresponding to each entity object, and when the semantic correlation is determined to be greater than a preset threshold, the input data is assigned to the matching result set corresponding to that entity object.

5. The method as described in claim 3 or 4, for each matching result set corresponding to an entity object, if the matching model determines that there are two different input data corresponding to any specified attribute of the entity object in the matching result set, then the input data with the highest confidence level of the corresponding data source is selected from the input data corresponding to the specified attribute as the standard input data corresponding to the specified attribute; the specified attribute is used to characterize the features or state of the entity object.

6. The method as described in claim 1, wherein the input data is input into a preset matching model, so that the matching model identifies input data belonging to the same entity object from the input data and merges them into the same matching result set, to obtain a matching result set output after this round of data matching, specifically including: The input data are fed into a preset matching model, which identifies input data belonging to the same entity object and merges them into the same matching result set to obtain a basic matching result set for each entity object. For each entity object, a profile feature data is generated based on the basic matching result set. Based on the profile feature data and the basic matching result set, a matching result set output after this round of data matching is obtained. The profile feature data is a derived feature of the entity object obtained by feature extraction or feature prediction on the basic matching result set.

7. The method as described in claim 1, wherein the input data is input into a preset matching model, so that the matching model identifies input data belonging to the same entity object from the input data and merges them into the same matching result set, specifically including: Based on the matching conditions of this round of data matching, at least one matching model is selected from the matching models contained in the preset model pool as the target matching model; the matching conditions include at least one of the following: the data type of the target data, the data source type to which the target data belongs, and the preset application scenario type; The input data is input into the target matching model so that the target matching model can identify input data belonging to the same entity object from each input data and merge them into the same matching result set.

8. The method as described in claim 1, wherein the input data is input into a preset matching model, so that the matching model identifies input data belonging to the same entity object from the input data and merges them into the same matching result set, specifically including: Select a target model group from the matching models included in the model pool, wherein the target model group contains at least two matching models; The input data is fed into a preset matching model so that the matching model identifies input data belonging to the same entity object and merges them into the same matching result set. Specifically, this includes: Each input data is input into each target matching model contained in the target matching model group. For each target matching model, the input data belonging to the same entity object are identified from each input data and merged into the same matching result set.

9. The method of claim 8, further comprising: Based on the weight coefficients of each target matching model in the target matching model group under the matching conditions of this round of data matching, the matching result sets output by each target matching model are fused to obtain the matching result set output by this round of data matching; wherein, the weight coefficient of any matching model under different matching conditions is determined based on the accuracy of the historical matching result set output by the matching model under the matching conditions.

10. The method as described in claim 1, wherein each matching result set includes the matching result sets output in different rounds of data matching; for each round of data matching, each matching result set output after that round of data matching includes a matching result set output by different matching models or combinations of different matching models contained in a preset model pool.

11. A business execution method, the method comprising: Receive data retrieval requests; The data retrieval request includes filtering conditions, which include at least one of the following: specifying a credibility level, specifying a matching model type, and specifying a business scenario. The target matching result set that matches the filtering criteria is determined from each pre-processed matching result set and returned.

12. The method of claim 11, wherein the target matching result set is sent to a target trusted space domain for storage, so that downstream business systems can call it in the target trusted space domain; the target trusted space domain is determined from preset trusted space domains based on the identifier of the data caller contained in the data call request; the trusted space domain is a logically isolated data storage and access space created for different data callers or groups of data callers.

13. A data matching device, comprising: The receiving module is used to receive the data to be processed sent by each data source; The matching module is used to perform at least two rounds of data matching on each piece of data to be processed to obtain matching result sets for use by downstream business systems. Each matching result set is used to represent a data profile of an entity object. In any round of data matching: based on the preset credibility level corresponding to each data source, the target data required for this round of data matching is determined from each piece of data to be processed. The higher the round of data matching, the higher the credibility level of the data source corresponding to the determined target data. The target data and the matched data contained in the matching result set output by the previous round of data matching are used as input data. The input data are input into a preset matching model so that the matching model can identify input data belonging to the same entity object from the input data and merge them into the same matching result set to obtain the matching result set output by this round of data matching.

14. An electronic device, comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method as described in any one of claims 1-12 by executing the executable instructions.

15. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-12.

16. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method as claimed in any one of claims 1-12.

Citation Information

Patent Citations

  • Data exporting method and device, electronic equipment and computer readable medium

    CN116881881A

  • Data acquisition method and device, electronic equipment and storage medium

    CN119621726A

  • Multi-person real-time risk assessment control system and method based on big data

    CN120634746A

  • Data quality detection method and device, medium and electronic equipment

    CN121255794A

  • Large language model-assisted entity name resolution

    US12511322B1