Business data recommendation method and device, computer equipment, readable storage medium and program product

By using blockchain dynamic certificate verification and hierarchical encryption technology, combined with reversible hash chains and dynamic decision trees, the data source screening process is optimized, solving the security and efficiency issues in data sharing, and realizing efficient and secure data recommendation and evaluation.

CN120874131APending Publication Date: 2025-10-31IND CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510746986.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies suffer from inefficient data source screening, and rigid security policies lead to data sharing risks and wasted computing resources, which particularly affect model training performance in federated learning.

Method used

The system verifies identity legitimacy by receiving a blockchain dynamic certificate, processes and distributes data with hierarchical encryption, classifies encrypted feature data to obtain category labels, generates business data recommendation results using a pre-set recommendation model, and optimizes the data processing flow by combining a reversible hash chain and a dynamic decision tree.

Benefits of technology

It significantly improves data security and evaluation efficiency, reduces the risk of key leakage and unauthorized access, optimizes encryption strategies and computing resource utilization, and enables automatic evaluation and recommendation of data source quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120874131A_ABST
    Figure CN120874131A_ABST
Patent Text Reader

Abstract

The invention relates to a business data recommendation method and device, computer equipment, a computer readable storage medium and a computer program product, and the method comprises the steps: receiving a participation request of each participant, and enabling the participation request to comprise a block chain dynamic certificate of each participant; if the block chain dynamic certificate is verified to be legal, performing hierarchical encryption processing on preset sample data to obtain encrypted data, and distributing the encrypted data to each participant; receiving encrypted feature data uploaded by each participant, and classifying the encrypted feature data to obtain a category label of the encrypted feature data, the encrypted feature data being data generated and encrypted after each participant matches local service data with the encrypted data; and generating a business data recommendation result according to the encrypted feature data, the category label and a preset recommendation model. According to the invention, the data security and the data source quality evaluation efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a business data recommendation method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology

[0002] The development of internet technology and artificial intelligence has greatly promoted the development of big data technology. Big data has extremely high value, and the need for data sharing and breaking down "data silos" is becoming increasingly prominent. However, data sharing carries data security risks such as privacy leaks, and against this backdrop, privacy computing technology has emerged.

[0003] Traditional privacy-preserving computing techniques, such as federated learning, require joint training of models without sharing the original data. The quality of the data source has a significant impact on the model training effect, so it is particularly important to select high-quality data sources before joint training of the models.

[0004] There are still some problems with data source selection, such as rigid security policies and the need for manual aggregation and analysis of data evaluation indicators, which leads to low efficiency. Summary of the Invention

[0005] Therefore, it is necessary to provide a business data recommendation method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve security and evaluation efficiency in response to the above-mentioned technical problems.

[0006] Firstly, this application provides a business data recommendation method, including:

[0007] Receive participation requests from each participant, wherein the participation requests include the blockchain dynamic certificates of each participant;

[0008] If the blockchain dynamic certificate is verified to be legitimate, the preset sample data will be encrypted in a hierarchical manner to obtain encrypted data, and the encrypted data will be distributed to each participant.

[0009] The system receives encrypted feature data uploaded by each participant, classifies the encrypted feature data, and obtains category labels for the encrypted feature data. The encrypted feature data is data generated and encrypted by each participant after matching local business data with the encrypted data.

[0010] Based on the encrypted feature data, the category labels, and the preset recommendation model, business data recommendation results are generated.

[0011] In one embodiment, if the blockchain dynamic certificate is verified to be legitimate, the preset sample data is subjected to hierarchical encryption processing to obtain encrypted data, including:

[0012] Based on a preset global field semantic graph, the field types of the sample data are obtained;

[0013] The sensitivity level of the sample data is determined based on the field type.

[0014] Obtain the corresponding encryption strategy based on the sensitivity level;

[0015] The sample data is encrypted according to the encryption strategy to obtain encrypted data.

[0016] In one embodiment, after verifying the validity of the blockchain dynamic certificate, the preset sample data is subjected to hierarchical encryption processing to obtain encrypted data, and the process further includes:

[0017] The hash value and hash alias of the encrypted data are generated using reversible hash chain technology, and the reversible mapping relationship between the hash alias, the hash value, and the encrypted data is stored in a preset blockchain network.

[0018] In one embodiment, the encrypted feature data is classified to obtain category labels for the encrypted feature data, including:

[0019] The encrypted feature data is semantically parsed using a pre-trained language model to obtain its semantic information.

[0020] The semantic information is input into a preset dynamic decision tree to obtain the category label of the encrypted feature data.

[0021] In one embodiment, the business data recommendation method further includes:

[0022] If the category label of the encrypted feature data is not obtained, the encrypted feature data and the semantic information are added to the clustering pool.

[0023] When the data in the data pool to be clustered reaches the preset conditions, clustering analysis of the data in the data pool to be clustered is triggered to obtain new labels;

[0024] The pre-trained language model is incrementally trained based on the incremental learning mechanism and the newly added labels;

[0025] The dynamic decision tree is updated using the incrementally trained language model;

[0026] The updated dynamic decision tree is then used as the preset dynamic decision tree, and the step of inputting the semantic information into the preset dynamic decision tree to obtain the category label of the encrypted feature data is returned.

[0027] In one embodiment, before reusing the updated dynamic decision tree as the preset dynamic decision tree, the method further includes:

[0028] If it is determined that there are similar labels in the dynamic decision tree, then the similar labels and their corresponding classification rules are marked as confused samples and stored in the dispute pool;

[0029] The obfuscated samples in the dispute pool are pushed to the review terminal for manual review, and the manual review results are obtained.

[0030] Based on the results of the manual review, the dynamic decision tree is updated using a reinforcement learning algorithm.

[0031] In one embodiment, a business data recommendation result is generated based on the encrypted feature data, the category label, and a preset recommendation model, including:

[0032] The encrypted feature data and the category label are input into the preset recommendation model;

[0033] The encrypted feature data is evaluated using the preset recommendation model to generate indicator data under different category labels. The indicator data under different category labels are weighted and calculated to obtain the comprehensive indicator data of each participant. Based on the comprehensive indicator data and preset business constraint rules, the business data recommendation result is output.

[0034] In one embodiment, before generating the business data recommendation result based on the encrypted feature data, the category label, and the preset recommendation model, the method further includes:

[0035] Obtain a training sample set, which includes indicator data under the type tags corresponding to the business data of the participants, as well as recommendation tags;

[0036] The initial recommendation model is trained based on the training sample set until it converges, thus obtaining the preset recommendation model.

[0037] Secondly, this application also provides a business data recommendation device, comprising:

[0038] A receiving module is used to receive participation requests from each participant, wherein the participation request includes the blockchain dynamic certificate of each participant;

[0039] An encryption module is used to perform hierarchical encryption processing on preset sample data to obtain encrypted data if the blockchain dynamic certificate is verified to be legitimate, and then distribute the encrypted data to each participant.

[0040] The classification module is used to receive the encrypted feature data uploaded by each participant, classify the encrypted feature data, and obtain the category label of the encrypted feature data. The encrypted feature data is obtained by each participant matching local business data with the encrypted data and encrypting it.

[0041] The result generation module is used to generate business data recommendation results based on the encrypted feature data, the category labels, and the preset recommendation model.

[0042] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the business data recommendation method in any of the above embodiments.

[0043] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the business data recommendation method in any of the above embodiments.

[0044] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the business data recommendation method in any of the above embodiments.

[0045] The aforementioned business data recommendation method, apparatus, computer equipment, computer-readable storage medium, and computer program product receive participation requests from various participants, including the blockchain dynamic certificates of each participant; if the blockchain dynamic certificates are verified to be valid, the preset sample data is hierarchically encrypted to obtain encrypted data, which is then distributed to each participant; encrypted feature data uploaded by each participant is received and classified to obtain category labels for the encrypted feature data, which is obtained by each participant matching and encrypting their local business data; and business data recommendation results are generated based on the encrypted feature data, category labels, and a preset recommendation model. Among these features, the use of blockchain dynamic certificates to verify the legitimacy of identities significantly reduces the risk of key leakage and unauthorized access, and improves data security, as the dynamic updates of blockchain dynamic certificates prevent brute-force attacks on static, long-term valid certificates. When the blockchain dynamic certificate is valid, pre-defined sample data is hierarchically encrypted before being distributed to all participants, avoiding rigid encryption strategies and wasted computing resources, thus improving data processing efficiency. Furthermore, by automatically classifying encrypted feature data to obtain category labels, and generating business data recommendation results based on encrypted feature data, category labels, and a pre-defined recommendation model, automatic evaluation of data source quality is achieved, significantly improving evaluation efficiency. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0047] Figure 1 This is a diagram illustrating the application environment of a business data recommendation method in one embodiment.

[0048] Figure 2 This is a flowchart illustrating a business data recommendation method in one embodiment;

[0049] Figure 3 This is a flowchart illustrating the specific steps involved in verifying the legitimacy of the blockchain dynamic certificate in one embodiment, namely, performing hierarchical encryption on preset sample data to obtain encrypted data.

[0050] Figure 4 This is a flowchart illustrating the steps that the business data recommendation method may further include in one embodiment;

[0051] Figure 5 This is a flowchart illustrating the steps included in one embodiment before the updated dynamic decision tree is used as the preset dynamic decision tree again.

[0052] Figure 6 This is a flowchart illustrating the specific steps included in S208 in one embodiment;

[0053] Figure 7 This is a structural block diagram of a business data recommendation device in one embodiment;

[0054] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0056] The business data recommendation method provided in this application embodiment can be applied to, for example, Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located in the cloud or on other network servers. Server 104 is the server for the participant who wants to lead the training process of the built-in model. The built-in model can be, for example, an anti-fraud model. Terminal 102 can be the terminal of other participants providing business data. These participants need to send a participation request before they can participate in model training, and their data source quality must be evaluated by server 104. Server 104 receives participation requests from terminals 102 of each participant. The participation requests include the blockchain dynamic certificates of each participant. After verifying the legality of the blockchain dynamic certificates, the server performs hierarchical encryption processing on the preset sample data to obtain encrypted data, and distributes the encrypted data to terminals 102. Server 104 receives encrypted feature data uploaded by terminals 102 and classifies the encrypted feature data to obtain category labels for the encrypted feature data. The encrypted feature data is obtained by terminals 102 matching local business data with the encrypted data and then encrypting it. Based on the encrypted feature data, category labels, and a preset recommendation model, server 104 generates business data recommendation results to guide participants who wish to lead the built-in model training process to select suitable business data providers to join the training.

[0057] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, etc. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0058] In one exemplary embodiment, such as Figure 2 As shown, a business data recommendation method is provided, which is applied to... Figure 1 Taking the server in the example, the explanation includes the following steps S202 to S208. Wherein:

[0059] S202, Receive participation requests from each participant, wherein the participation requests include the blockchain dynamic certificates of each participant.

[0060] The server receives participation requests from the terminals of each participating party, and these requests include the participating party's blockchain dynamic certificate. A blockchain dynamic certificate is a digital identity credential issued based on blockchain technology, supporting automatic updates and tamper-proof verification. Examples may include, but are not limited to, a bank's digital certificate chain or an insurance company's digital certificate chain.

[0061] S204. If the blockchain dynamic certificate is verified to be legitimate, the preset sample data will be encrypted in a hierarchical manner to obtain encrypted data, and the encrypted data will be distributed to each participant.

[0062] After receiving participation requests from various parties, the server needs to verify the legitimacy of the blockchain dynamic certificates, including verifying their validity and scope of authorization. Traditional identity authentication primarily relies on static API key authentication, which carries the security risk of unauthorized access due to key leakage. Blockchain dynamic certificates, however, can be dynamically updated, preventing long-term valid certificates from being brute-forced, thus reducing the risk of key leakage and unauthorized access, and improving data security. In one example, the server can verify the legitimacy of the blockchain dynamic certificate through a blockchain network. This blockchain network establishes a connection with the server and is used for identity verification and log auditing, improving the traceability of data access. For instance, if the authorization scope of a P2P lending platform's blockchain dynamic certificate does not include the user's contact list data, and the platform frequently requests contact list data, the blockchain network will verify that the permission request exceeds the authorized scope, automatically blocking it, triggering an alarm, and recording the permission request log.

[0063] When verifying the legitimacy of the blockchain dynamic certificate, independent Docker sandboxes can be dynamically created and allocated to each participant to confine all running applications and processes within that sandbox. For example, the pre-defined sample data hierarchical encryption processing and distribution of encrypted data, as well as subsequent steps S206 and S208, all need to be performed within each participant's respective sandbox.

[0064] Sandboxes can also be enhanced. For example, eBPF can be used to monitor container network traffic and detect abnormal data transmission; Seccomp-BPF policies can be used to restrict system call permissions for processes within the sandbox. eBPF is a Linux kernel extension technology used for real-time monitoring and filtering of container network traffic and other system behaviors; Seccomp-BPF is a Linux security module that prevents malicious operations by restricting process system calls.

[0065] By using blockchain dynamic certificates and sandbox isolation, the risk of unauthorized access can be further reduced, and data security can be improved.

[0066] The preset sample data consists of user data pre-stored on the server, such as personal ID numbers, contact information, addresses, and monthly salaries. The server performs hierarchical encryption on the preset sample data to obtain encrypted data. This addresses the problems of rigid encryption strategies in traditional technologies and the waste of computational resources for low-sensitivity fields caused by uniform encryption strategies, thus balancing encryption security with the rational allocation of computational resources.

[0067] It is understood that when the server distributes the encrypted data to each participant, it is to distribute the encrypted data within the authorization scope of each participant's blockchain dynamic certificate to the corresponding participant.

[0068] S206, receive the encrypted feature data uploaded by each participant, classify the encrypted feature data, and obtain the category label of the encrypted feature data. The encrypted feature data is obtained by each participant matching local business data with the encrypted data and encrypting it.

[0069] In this step, the server receives and categorizes the encrypted feature data returned by each participant. In an optional instance, this step is executed within the sandbox corresponding to each participant. Each participant, upon receiving the encrypted data distributed by the server, needs to match its local business data with the received encrypted data. For example, it might use the last four digits of the ID number in the encrypted data to find the user's repayment data under that ID number, and then further score or statistically analyze the user's repayment data according to the participant's local preset rules, forming repayment score data or statistical data. These preset rules could be on-time repayment status, overdue status, etc. The matching process refers to the process of reprocessing or calculating the business data matching the encrypted data according to certain preset rules to generate processed data.

[0070] In one example, the participants encrypt the generated processing data again before uploading it to the server. The encryption process includes, but is not limited to, using reversible encryption algorithms and / or differential privacy (DP) processing.

[0071] In one implementation, the server classifies the encrypted feature data to obtain category labels for the encrypted feature data. Specifically, this includes: performing semantic parsing on the encrypted feature data using a pre-trained language model to obtain semantic information of the encrypted feature data; and inputting the semantic information into a preset dynamic decision tree to obtain category labels for the encrypted feature data.

[0072] In this implementation, the pre-trained language model, such as the BERT model, can obtain corresponding semantic information after semantic parsing of the encrypted feature data. For example, the encrypted feature data is the de-identified user's repayment score "User ****0001 has been overdue once in the past 6 months, with a score of 95". The BERT model is input to obtain a semantic vector, which means "low number of overdue payments in the past 6 months, excellent score". It can be understood that the BERT model may need to decrypt the encrypted feature data first to obtain input data that can be understood by the model, or it may directly parse the encrypted feature data that has only been de-identified, as long as the encrypted feature data can be understood by the BERT model.

[0073] The server takes semantic information, such as "low number of overdue payments in the past 6 months and excellent credit score", and inputs it into a preset dynamic decision tree after transformation such as dimensionality reduction, labeling or extraction of semantic keywords. This will obtain the category label of the encrypted feature data. In this example, the category label that can be obtained can be a credit label, such as "good credit".

[0074] S208, Based on the encrypted feature data, the category label, and the preset recommendation model, generate business data recommendation results.

[0075] Specifically, the server inputs the encrypted feature data and category labels into the recommendation model, and the model processes the data to obtain the output result, which serves as the business data recommendation result.

[0076] For example, the business data recommendation results can be of various data types such as text, images, lists, audio, or video, or a combination of multiple data types. For instance, the business data recommendation results can be a recommendation report, with content presented in a graphic and textual format. Specific content may include evaluation metrics for each participant's business data, metric values, evaluation results regarding their suitability for training specific models, suggestions for combining participant business data, and so on.

[0077] Understandably, the parties who wish to lead the built-in model training process can select suitable parties that provide business data to join the model training process based on the above business data recommendation results, and use known techniques such as federated learning and secure multi-party computation to jointly train the target model, which will not be described in detail in this article.

[0078] The business data recommendation method in this embodiment receives participation requests from various participants, including their blockchain dynamic certificates. If the blockchain dynamic certificates are verified to be valid, preset sample data is hierarchically encrypted to obtain encrypted data, which is then distributed to each participant. The method also receives encrypted feature data uploaded by each participant, classifies the encrypted feature data to obtain category tags, and obtains the encrypted feature data by matching and encrypting local business data. Finally, based on the encrypted feature data, category tags, and a preset recommendation model, a business data recommendation result is generated. Among these features, the use of blockchain dynamic certificates to verify the legitimacy of identities significantly reduces the risk of key leakage and unauthorized access, and improves data security, as the dynamic updates of blockchain dynamic certificates prevent brute-force attacks on static, long-term valid certificates. When the blockchain dynamic certificate is valid, pre-defined sample data is hierarchically encrypted before being distributed to all participants, avoiding rigid encryption strategies and wasted computing resources, thus improving data processing efficiency. Furthermore, by automatically classifying encrypted feature data to obtain category labels, and generating business data recommendation results based on encrypted feature data, category labels, and a pre-defined recommendation model, automatic evaluation of data source quality is achieved, significantly improving evaluation efficiency.

[0079] In one exemplary embodiment, such as Figure 3 As shown, if the blockchain dynamic certificate is verified to be legitimate, the preset sample data is subjected to hierarchical encryption processing to obtain encrypted data, including steps S302 to S308. Wherein:

[0080] S302, Based on the preset global field semantic graph, obtain the field types of the sample data.

[0081] The global field semantic graph is a pre-built semantic network based on knowledge graph technology. It aims to align the differentiated representations of the same field by different participants, eliminate field ambiguity, and achieve structured processing. By standardizing semantics and associating heterogeneous fields, the global field semantic graph quickly aligns field meanings in indicator mapping scenarios, reducing the cost of manual mapping.

[0082] For example, participant A uses "monthly income" to refer to total after-tax income, while participant B uses "pre-tax monthly salary" to refer to total pre-tax income. The global field semantic graph will create nodes for these two fields, label their specific definitions (such as "pre-tax monthly salary = pre-tax monthly income", "monthly income = after-tax income"), and associate them with standard concepts (such as "personal income").

[0083] S304, Determine the sensitivity level of the sample data based on the field type.

[0084] There is a mapping relationship between field types and sensitivity levels; different field types correspond to different sensitivity levels. Once the field types of the sample data are determined, the sensitivity level of the sample data can be further determined based on the field types. For example, sensitivity levels can be divided into "high sensitivity," "medium sensitivity," and "low sensitivity." In one example, the ID number field corresponds to high sensitivity, the user address field and the consumption amount field correspond to medium sensitivity, and the user name and age fields correspond to low sensitivity.

[0085] S306, Obtain the corresponding encryption strategy based on the sensitivity level.

[0086] In this step, the server further obtains the corresponding encryption strategy based on the sensitivity level of the sample data. Understandably, a mapping relationship is established between sensitivity level and encryption strategy. For example, high sensitivity → SM9 identifier cryptography algorithm, medium sensitivity → AES symmetric encryption algorithm, low sensitivity → lightweight encryption algorithm (such as ChaCha20).

[0087] S308, The sample data is encrypted according to the encryption strategy to obtain encrypted data.

[0088] For example, an encryption engine middleware can also be designed in the server. Once the encryption strategy to be used for the sample data is determined, the server drives the encryption engine middleware to dynamically load algorithms such as SM9 / AES / ChaCha20 to encrypt the sample data, thus obtaining encrypted data. For example, "ID number" → high-sensitivity → SM9 encryption, "consumption amount" → medium-sensitivity → AES encryption.

[0089] In this embodiment, the field types of sample data are obtained based on a preset global field semantic graph, which can quickly align the meaning of fields in the indicator mapping scenario, reduce the cost of manual mapping, and improve the efficiency and accuracy of data processing. The sensitivity level of the sample data is determined according to the field type, and then the corresponding encryption strategy is obtained according to the sensitivity level. The sample data is then encrypted according to the encryption strategy, which optimizes the utilization rate of encryption resources and further solves the problems of rigid encryption strategies in traditional technology and waste of computing resources for low-sensitivity fields caused by unified encryption strategies. It balances encryption security and the rationality of computing resource allocation, and also helps to improve global computing efficiency.

[0090] In one embodiment, the method further includes the following step after step S204:

[0091] S205, generate the hash value and hash alias of the encrypted data through reversible hash chain technology, and store the hash alias, the hash value and the reversible mapping relationship of the encrypted data in a preset blockchain network.

[0092] Reversible hash chain technology refers to a technique that supports the reverse tracing of the original field mapping of encrypted data through hash aliases. Traditional hash functions are unidirectional and irreversible. The reversible hash chain technology described in this paper combines encryption and hashing techniques. Specifically, it calculates a hash value from encrypted data, generates a unique hash alias for the hash value, hides the original hash information, and then stores the original encrypted data, hash value, and hash alias together in a blockchain network. Taking a user address field as an example, the server first determines that the user address field is "medium-sensitive" data, then determines an intermediate encryption algorithm (such as AES), generates an encrypted formatted address, creates a unique hash value for the encrypted formatted address, generates a unique hash alias based on this hash value, and stores the encrypted formatted address, hash value, and hash alias together in a secure environment (such as a blockchain network). When it is necessary to calculate or audit the user address, the hash alias and hash value are used to restore the encrypted formatted address, and the original user address is obtained using the key.

[0093] This embodiment combines sensitivity-level encryption technology with reversible hash chain technology to support subsequent data traceability and integrity verification, which can shorten the audit location time and enhance data traceability.

[0094] The above description of S206 specifically includes: performing semantic parsing on encrypted feature data using a pre-trained language model to obtain semantic information of the encrypted feature data; and inputting the semantic information into a preset dynamic decision tree to obtain the category label of the encrypted feature data. However, not all encrypted feature data can find a corresponding category label through a preset dynamic decision tree after obtaining semantic information through a pre-trained language model. Especially when new business data appears, the existing branches of the dynamic decision tree may not be able to cover the new business scenario. How to improve the classification accuracy in complex scenarios and enhance the flexibility and business adaptability of the label system are also technical problems that need to be further solved.

[0095] Therefore, in one exemplary embodiment, such as Figure 4 As shown, the business data recommendation method further includes steps S402 to S410. Wherein:

[0096] S402, if the category label of the encrypted feature data is not obtained, the encrypted feature data and the semantic information are added to the clustering data pool.

[0097] The data pool to be clustered refers to the set of data that has not yet been classified or grouped. This data needs to be analyzed for similarity using clustering algorithms and divided into different clusters. For example, the encrypted feature data "User ****0001's total live stream reward amount in the past week was 5000 yuan, all paid by credit card" has semantic information in the form of a semantic vector after semantic parsing by a pre-trained language model. Its corresponding meaning is "User's live stream reward amount reached 5000 yuan and only paid by credit card". After this semantic vector is transformed and input into a preset dynamic decision tree, it does not obtain a corresponding category label. Therefore, the encrypted feature data "User ****0001's total live stream reward amount in the past week was 5000 yuan, all paid by credit card" and the corresponding semantic information are added to the data pool to be clustered.

[0098] S404, when the data in the data pool to be clustered reaches the preset conditions, clustering analysis of the data in the data pool to be clustered is triggered to obtain new labels.

[0099] Understandably, the preset conditions include, but are not limited to, the amount of unprocessed data in the data pool exceeding a preset threshold, new data remaining unprocessed for a certain time window, the semantic vector distribution of new data differing significantly from historical data, the triggering of scheduled tasks, the triggering of computing resource availability, or the update of the clustering model version, etc.

[0100] When preset conditions are met, the server triggers cluster analysis on the data in the data pool to be clustered, and obtains new labels. Taking live streaming reward data as an example, when the data in the data pool to be clustered includes multiple data points about live streaming reward situations, the data that meet the similarity conditions are divided into a set of datasets through cluster analysis, and their new labels are determined, such as "large amount live streaming reward transaction".

[0101] S406, Incrementally train the pre-trained language model based on the incremental learning mechanism and the newly added labels.

[0102] The incremental learning mechanism is a training mechanism that allows the model to continuously learn new data without forgetting old knowledge. In this step, the server incrementally trains the pre-trained language model based on the incremental learning mechanism and the newly added labels. For example, an incremental training set is constructed based on samples with newly added labels and original historical data, and the parameters of the BERT model are fine-tuned in combination with a preset training strategy to obtain the incrementally trained language model.

[0103] S408, Update the dynamic decision tree using the incrementally trained language model.

[0104] In one example, the server uses the incrementally trained language model to generate new business rules. After processing the new business rules with feature engineering, a high-dimensional input to the decision tree is formed, which in turn forms new branch conditions for the dynamic decision tree. For example, the new branch condition is "if the reward amount is ≥ 5000 yuan", and the new node is "high-value user".

[0105] S410, the updated dynamic decision tree is used again as the preset dynamic decision tree, and the step of inputting the semantic information into the preset dynamic decision tree to obtain the category label of the encrypted feature data is returned.

[0106] In this step, the updated dynamic decision tree is used to process new business scenarios to meet the data classification requirements in these scenarios. For example, the updated dynamic decision tree is used again as the preset dynamic decision tree. When encrypted feature data related to live streaming rewards is input into the language model (i.e., the incrementally trained language model), and the semantic information obtained is that the reward amount is 6,000 yuan, the dynamic decision tree can determine the category label corresponding to the encrypted feature data as a high-value user based on this semantic information.

[0107] In this embodiment, steps S402 to S410 enable continuous and automated updates of the preset language model and dynamic decision tree, improving the system's adaptability. This allows the system to better adapt to business changes, improve classification accuracy in complex scenarios, and enhance the flexibility and adaptability of the labeling system.

[0108] In one exemplary embodiment, such as Figure 5 As shown, before the updated dynamic decision tree is used again as the preset dynamic decision tree, steps S502 to S506 are included. Wherein,

[0109] S502, if it is determined that there are similar labels in the dynamic decision tree, then the similar labels and the corresponding unclassified samples are marked as confused samples and stored in the dispute pool.

[0110] For example, when certain samples to be classified are input into a pre-trained language model for parsing and the category labels are determined by a pre-set dynamic decision tree, two category labels are obtained, such as label A: income stability score with a confidence level of 0.65 and label B: comprehensive income score with a confidence level of 0.68. Since the confidence levels of both labels are less than the confidence threshold (0.7), the samples to be classified, as well as the obtained labels A, B, confidence levels, and classification paths in the decision tree, are all marked as confused samples and stored in the dispute pool.

[0111] S504, push the obfuscated samples in the dispute pool to the review terminal for manual review and obtain the manual review results.

[0112] In this step, the manual review result can be the correct category label for the confused sample, and the manual review result is also stored in the dispute pool again. For example, if the manual review determines that the correct category label for the sample to be classified is label B: comprehensive income score, then label B is recorded as the manual review result and further stored in the corresponding confused sample data in the dispute pool.

[0113] S506, Based on the results of the manual review, update the dynamic decision tree using a reinforcement learning algorithm.

[0114] Among them, reinforcement learning algorithms are machine learning methods that optimize decision-making strategies through trial and error feedback, and are used for handling controversial samples. Specifically, updating the dynamic decision tree using reinforcement learning algorithms involves using manually reviewed controversial samples as "reward signal" samples for reinforcement learning. For example, when label B is the correct category label, its branch rule is given a positive reward of +1, and when label A is the incorrect category label, its branch rule is given a negative reward of -0.5.

[0115] In this embodiment, by acquiring confused samples and introducing manual review and reinforcement learning algorithms, a confused label arbitration mechanism is further introduced for the evolution of dynamic decision trees, thereby optimizing the structure of the decision tree or the label mapping rules, further reducing classification bias caused by semantic ambiguity, and improving classification accuracy.

[0116] In one embodiment, such as Figure 6 As shown, step S208 above includes steps S602 and S604. Wherein:

[0117] S602, input the encrypted feature data and the category label into the preset recommendation model.

[0118] S604, the preset recommendation model calculates the indicators of the encrypted feature data to generate indicator data under different category labels, and calculates the weighted indicator data under different category labels to obtain the comprehensive indicator data of each participant. Based on the comprehensive indicator data and preset business constraint rules, the business data recommendation result is output.

[0119] Specifically, the indicator data is used to evaluate the quality of business data under different category labels and its impact on the target built-in model. It mainly includes: PSI value (Population Stability Index), an indicator that measures the stability of data distribution and is used to detect data shift. The smaller the value, the more stable the data distribution; KS value (Kolmogorov-Smirnov), an indicator that evaluates the model's discrimination ability and reflects the difference between positive and negative sample distributions. The larger the value, the stronger the model's discrimination ability; AUC value (Area Under Curve), an indicator that evaluates the overall performance of the classification model. The closer the value is to 1, the better the performance.

[0120] The pre-defined recommendation model calculates encrypted feature data metrics for different category labels, obtaining metric data for each category label. For each category label, KS, PSI, and AUC are obtained. The metric data for different category labels are then weighted and calculated, specifically according to metric type, to obtain the comprehensive metric data for each participant under various metric types. The recommendation model then outputs business data recommendation results based on the comprehensive metric data and pre-defined business rule constraints. These pre-defined business rule constraints can be set according to the requirements of the target built-in model for business data.

[0121] Taking the built-in anti-fraud model as an example, the following explanation is provided. Participant A, providing business data, has encrypted feature data corresponding to category labels a, b, and c. The server inputs label a and its encrypted feature data, label b and its encrypted feature data, and label c and its encrypted feature data into a pre-defined recommendation model for indicator calculation, obtaining indicator data for each category label. In the anti-fraud scenario, assuming label a has a weight of 30% (medium business importance), label b has a weight of 50% (highest business importance), and label c has a weight of 20% (lowest business importance), the indicator data under different category labels are weighted and calculated to obtain comprehensive indicator data, as shown in Table 1.

[0122] Table 1. Indicator data and comprehensive indicator data under each category label of participant A.

[0123]

[0124] In other words, the final composite KS of participant A is 0.45, composite AUC is 0.81, and composite PSI is 0.10.

[0125] If there are three participating parties, A, B, and C, their comprehensive indicator data are shown in Table 2:

[0126] Table 2. Comprehensive indicator data for each participating party

[0127]

[0128] For example, business rule constraints for anti-fraud scenarios could be: "KS≥0.3 and PSI≤0.15, if PSI>0.15, mandatory elimination".

[0129] The business data recommendation results output by the recommendation model based on the above comprehensive indicator data and preset business rule constraints can be a recommendation report, which can include the summary of indicator data of each participant, such as the comprehensive indicators of each participant, indicators under different types of labels, and whether each participant can participate in the training of the target built-in model, and suggestions for the combined application of business data of each participant.

[0130] Continuing with the example of participants A, B, and C, the core content of the recommendation report output by the recommendation model could include, for example: "Participant A, recommended, can be used as a core data source for training the anti-fraud model; Participant B, forced elimination, PSI value fluctuates abnormally (0.18); Participant C, recommended elimination, KS does not meet the standard," and could also include: "It is recommended to prioritize the use of participant A's data source, and participant C and participant A can be used together in the short term, but KS needs to be optimized to improve the discrimination ability," and so on.

[0131] In this embodiment, encrypted feature data and category labels are input into a preset recommendation model. The recommendation model calculates various indicators on the encrypted feature data to obtain comprehensive indicator data for each participant. Then, based on the comprehensive indicator data and preset business constraint rules, it outputs business data recommendation results. This realizes automatic evaluation and summarization of data source quality and provides recommendation suggestions, which significantly improves evaluation efficiency and provides strong support for decision-makers to select high-quality data sources and obtain efficient and high-quality training results during subsequent model training.

[0132] In an exemplary embodiment, steps S101 and S103 are included before step S208. Wherein:

[0133] S101, Obtain the training sample set, which includes indicator data under the type tags corresponding to the business data of the participants, comprehensive indicator data, and recommendation tags.

[0134] For example, the training sample set includes indicator data KS, PSI, and AUC under the type labels of each participant's business data, as well as the comprehensive KS, PSI, and AUC of each participant, and recommendation labels ("recommended" or "not recommended"). The training samples can be derived from historical business data, such as participant data used in the past when training the anti-fraud model.

[0135] S103, The initial recommendation model is trained based on the training sample set until the initial recommendation model converges, thus obtaining the preset recommendation model.

[0136] Specifically, the initial model can be, for example, a logistic regression model, a tree model (XGBoost), or a weighted linear combination model, etc. The initial model embeds business constraint rules through a preset loss function, such as hard constraints: PSI > 0.15 → not recommended; weighted calculation of weight allocation, etc. The specific model training process in this step can refer to a general model training workflow, inputting the training sample set into the initial recommendation model, and continuously iterating and optimizing it through training loops, validation and parameter tuning, testing and evaluation, until the initial recommendation model converges.

[0137] In this embodiment, the initial model is trained to obtain a preset recommendation model, which lays the foundation for the smooth implementation of the business data recommendation method.

[0138] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0139] Based on the same inventive concept, this application also provides a business data recommendation apparatus for implementing the business data recommendation method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more business data recommendation apparatus embodiments provided below can be found in the limitations of the business data recommendation method described above, and will not be repeated here.

[0140] In one exemplary embodiment, such as Figure 7 As shown, a business data recommendation device 700 is provided, including: a receiving module 702, an encryption module 704, a classification module 706, and a result generation module 708, wherein:

[0141] The receiving module 702 is used to receive participation requests from each participant, wherein the participation request includes the blockchain dynamic certificate of each participant;

[0142] The encryption module 704 is used to perform hierarchical encryption processing on the preset sample data to obtain encrypted data if the blockchain dynamic certificate is verified to be legitimate, and then distribute the encrypted data to each participant.

[0143] The classification module 706 is used to receive the encrypted feature data uploaded by each participant, classify the encrypted feature data, and obtain the category label of the encrypted feature data. The encrypted feature data is obtained by each participant matching local business data with the encrypted data and encrypting it.

[0144] The result generation module 708 is used to generate business data recommendation results based on the encrypted feature data, the category labels, and the preset recommendation model.

[0145] In one embodiment, the encryption module 704 is further configured to: obtain the field type of the sample data based on a preset global field semantic graph; determine the sensitivity level of the sample data according to the field type; obtain the corresponding encryption strategy according to the sensitivity level; and encrypt the sample data according to the encryption strategy to obtain encrypted data.

[0146] In one embodiment, the business data recommendation device 700 further includes:

[0147] The hash generation and storage module is used to generate the hash value and hash alias of the encrypted data through reversible hash chain technology, and store the hash alias, the hash value and the reversible mapping relationship of the encrypted data in a preset blockchain network.

[0148] In an exemplary embodiment, the classification module 706 is further configured to: perform semantic parsing on the encrypted feature data using a pre-trained language model to obtain semantic information of the encrypted feature data; and input the semantic information into a preset dynamic decision tree to obtain the category label of the encrypted feature data.

[0149] In one embodiment, the business data recommendation device 700 further includes:

[0150] The clustering analysis module is used to add the encrypted feature data and the semantic information to the data pool to be clustered if no category label is obtained for the encrypted feature data; when the data in the data pool to be clustered reaches a preset condition, the clustering analysis of the data in the data pool to be clustered is triggered to obtain new labels.

[0151] An incremental training module is used to incrementally train the pre-trained language model based on an incremental learning mechanism and the newly added labels.

[0152] The update module is used to update the dynamic decision tree using the incrementally trained language model, and to use the updated dynamic decision tree as the preset dynamic decision tree again.

[0153] In one embodiment, the business data recommendation device 700 further includes:

[0154] The confusion arbitration module is used to mark the similar labels and the corresponding unclassified samples as confused samples and store them in the dispute pool if it is determined that there are similar labels in the dynamic decision tree; and to push the confused samples in the dispute pool to the review terminal for manual review and obtain the manual review result.

[0155] The update module is also used to update the dynamic decision tree using a reinforcement learning algorithm based on the results of the manual review.

[0156] In one embodiment, the result generation module 708 is further configured to: input the encrypted feature data and the category label into the preset recommendation model; perform index calculation on the encrypted feature data by the preset recommendation model to generate index data under different category labels; perform weighted calculation on the index data under different category labels to obtain the comprehensive index data of each participant; and output the business data recommendation result according to the comprehensive index data and the preset business constraint rules.

[0157] In one embodiment, the business data recommendation device 700 further includes:

[0158] The recommendation model training module is used to obtain a training sample set, which includes indicator data under the type tags corresponding to the business data of the participants, as well as recommendation tags; and to train the initial recommendation model based on the training sample set until the initial recommendation model converges to obtain the preset recommendation model.

[0159] Each module in the aforementioned business data recommendation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0160] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores relevant task data for business data recommendation tasks. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a business data recommendation method.

[0161] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0162] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the business data recommendation method in any of the above embodiments.

[0163] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the business data recommendation method in any of the above embodiments.

[0164] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the business data recommendation method in any of the above embodiments.

[0165] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0166] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0167] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0168] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A business data recommendation method, characterized in that, The method includes: Receive participation requests from each participant, wherein the participation requests include the blockchain dynamic certificates of each participant; If the blockchain dynamic certificate is verified to be legitimate, the preset sample data will be encrypted in a hierarchical manner to obtain encrypted data, and the encrypted data will be distributed to each participant. The system receives encrypted feature data uploaded by each participating party, classifies the encrypted feature data, and obtains category labels for the encrypted feature data. The encrypted feature data is obtained by each participating party matching local business data with the encrypted data and then encrypting it. Based on the encrypted feature data, the category labels, and the preset recommendation model, business data recommendation results are generated.

2. The method according to claim 1, characterized in that, If the blockchain dynamic certificate is verified to be legitimate, the preset sample data will be subjected to hierarchical encryption processing to obtain encrypted data, including: Based on a preset global field semantic graph, the field types of the sample data are obtained; The sensitivity level of the sample data is determined based on the field type. Obtain the corresponding encryption strategy based on the sensitivity level; The sample data is encrypted according to the encryption strategy to obtain encrypted data.

3. The method according to claim 1, characterized in that, If the blockchain dynamic certificate is verified to be legitimate, then after performing hierarchical encryption on the preset sample data to obtain encrypted data, the method further includes: The hash value and hash alias of the encrypted data are generated using reversible hash chain technology, and the reversible mapping relationship between the hash alias, the hash value, and the encrypted data is stored in a preset blockchain network.

4. The method according to claim 1, characterized in that, The process of classifying the encrypted feature data to obtain category labels for the encrypted feature data includes: The encrypted feature data is semantically parsed using a pre-trained language model to obtain its semantic information. The semantic information is input into a preset dynamic decision tree to obtain the category label of the encrypted feature data.

5. The method according to claim 4, characterized in that, The method further includes: If the category label of the encrypted feature data is not obtained, the encrypted feature data and the semantic information are added to the clustering pool. When the data in the data pool to be clustered reaches the preset conditions, clustering analysis of the data in the data pool to be clustered is triggered to obtain new labels; The pre-trained language model is incrementally trained based on the incremental learning mechanism and the newly added labels; The dynamic decision tree is updated using the incrementally trained language model; The updated dynamic decision tree is then used as the preset dynamic decision tree, and the step of inputting the semantic information into the preset dynamic decision tree to obtain the category label of the encrypted feature data is returned.

6. The method according to claim 5, characterized in that, Before the updated dynamic decision tree is used again as the preset dynamic decision tree, the method further includes: If it is determined that there are similar labels in the dynamic decision tree, then the similar labels and the corresponding unclassified samples are marked as confused samples and stored in the dispute pool; The obfuscated samples in the dispute pool are pushed to the review terminal for manual review, and the manual review results are obtained. Based on the results of the manual review, the dynamic decision tree is updated using a reinforcement learning algorithm.

7. The method according to any one of claims 1 to 6, characterized in that, The step of generating business data recommendation results based on the encrypted feature data, the category labels, and the preset recommendation model includes: The encrypted feature data and the category label are input into the preset recommendation model; The encrypted feature data is evaluated using the preset recommendation model to generate indicator data under different category labels. The indicator data under different category labels are weighted and calculated to obtain the comprehensive indicator data of each participant. Based on the comprehensive indicator data and preset business constraint rules, the business data recommendation result is output.

8. The method according to any one of claims 1 to 6, characterized in that, Before generating business data recommendation results based on the encrypted feature data, the category labels, and the preset recommendation model, the method further includes: Obtain a training sample set, which includes indicator data under the type tags corresponding to the business data of the participants, as well as recommendation tags; The initial recommendation model is trained based on the training sample set until it converges, thus obtaining the preset recommendation model.

9. A business data recommendation device, characterized in that, The device includes: A receiving module is used to receive participation requests from each participant, wherein the participation request includes the blockchain dynamic certificate of each participant; An encryption module is used to, if the blockchain dynamic certificate is verified to be legitimate, perform hierarchical encryption processing on preset sample data to obtain encrypted data, and distribute the encrypted data to each participant. The classification module is used to receive the encrypted feature data uploaded by each participant, classify the encrypted feature data, and obtain the category label of the encrypted feature data. The encrypted feature data is obtained by each participant matching local business data with the encrypted data and encrypting it. The result generation module is used to generate business data recommendation results based on the encrypted feature data, the category labels, and the preset recommendation model.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.