A method and device for verifying derivative data, a terminal device, and a storage medium
By performing clustering spatial verification of the derived data of the ChatGPT model, the shortcomings of data reliability verification are solved, processing efficiency and accuracy are improved, and it is suitable for data analysis in finance and other fields.
Patent Information
- Application Number
- CN202311245531.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-09-25
AI Technical Summary
The prior art lacks reliability verification methods for data derived from ChatGPT models, resulting in unstable and unreliable in domain-specific information understanding and application.
By acquiring sample data, clustering the data of the first cluster space and the second cluster space using a preset clustering algorithm, the first and second clustering results are obtained, and the derivative data reliability of the preset language model output is verified based on these results.
The reliability of the derived data output from the preset language model is effectively verified, and the user processing efficiency is improved, especially in applications such as risk control analysis and customer ratings, no need to reconstruct the analysis model.
Smart Images

Figure CN117370554B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of data processing, and particularly relates to a method and device for verifying derivative data, a terminal device, and a storage medium. Background Art
[0002] With the development of technology, more and more chat tools based on language models are being used by people. For example, the currently popular ChatGPT (Chat Generative Pre-trained Transformer). ChatGPT is a natural language processing tool driven by artificial intelligence technology. It can generate answers based on the patterns and statistical laws seen during the pre-training stage and can also interact according to the context of the chat.
[0003] However, in various fields currently, there are rich domain-specific information. How to use the above ChatGPT-like models to stably, reliably, and effectively understand this domain knowledge and embed it into the vector space is an important and novel research topic currently. In other words, there is currently a lack of a way to verify the reliability of the data derived from ChatGPT-like models. Summary of the Invention
[0004] Embodiments of this application provide a method and device for verifying derivative data, a terminal device, and a storage medium, which can then effectively verify the reliability of the derivative data generated by a preset language model.
[0005] In a first aspect, embodiments of this application provide a method for verifying derivative data, including: obtaining sample data; where the sample data includes the original data and derivative data of multiple users; the derivative data is generated by a preset language model; the preset language model is used to analyze the input original data and generate the derivative data required by the user; clustering the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain a first clustering result and a second clustering result; where the first clustering space includes the original data of the multiple users; the second clustering space includes the derivative data of the multiple users; the first clustering result corresponds to the first clustering space; the second clustering result corresponds to the second clustering space; verifying the reliability of the derivative data output by the preset language model based on the first clustering result and the second clustering result.
[0006] In a possible implementation of the first aspect, clustering the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain a first clustering result and a second clustering result includes: performing clustering initialization on the first clustering space to obtain n clustering clusters, and generating n clustering clusters with the same categories in the second clustering space; n is a positive integer greater than 1; performing a clustering operation on the first clustering space to obtain the clustering result of the n clustering clusters in the first clustering space; randomly sampling each clustering cluster in the first clustering space to obtain a sampling result; using the sampling result as the initial clustering centers of the n clustering clusters in the second clustering space, and performing a clustering operation on the second clustering space to obtain the clustering result of the n clustering clusters in the second clustering space; wherein, the first clustering result includes the clustering result of the n clustering clusters in the first clustering space; the second clustering result includes the clustering result of the n clustering clusters in the second clustering space.
[0007] In a possible implementation of the first aspect, clustering the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain a first clustering result and a second clustering result includes: S1: performing clustering initialization on the first clustering space to obtain n clustering clusters, and generating n clustering clusters with the same categories in the second clustering space; n is a positive integer greater than 1; S2: performing a clustering operation on the first clustering space to obtain the clustering result of the n clustering clusters in the first clustering space; S3: randomly sampling each clustering cluster in the first clustering space to obtain the sampling result of the first clustering space; S4: using the sampling result of the first clustering space as the initial clustering centers of the n clustering clusters in the second clustering space, and performing a clustering operation on the second clustering space to obtain the clustering result of the n clustering clusters in the second clustering space; S5: randomly sampling each clustering cluster in the second clustering space to obtain the sampling result of the second clustering space; S6: using the sampling result of the second clustering space as the initial clustering centers of the n clustering clusters in the first clustering space, and performing a clustering operation on the first clustering space to obtain the clustering result of the n clustering clusters in the first clustering space; S7: repeating the above steps S3 - S6 until the clustering results of the n clustering clusters in the first clustering space and the clustering results of the n clustering clusters in the second clustering space converge; S8: outputting the first clustering result and the second clustering result; wherein, the first clustering result includes the clustering result of the n clustering clusters in the converged first clustering space; the second clustering result includes the clustering result of the n clustering clusters in the converged second clustering space.
[0008] In a possible implementation of the first aspect, after verifying the reliability of the derived data output by the preset language model, the method further includes: determining the weight values of the respective derived data; aggregating the original data of the multiple users, the derived data of the multiple users, and the weight values of the derived data of the multiple users to generate target data.
[0009] In a possible implementation of the first aspect, the determining the weight values of the respective derived data includes: obtaining the total number of users in the intersection of each cluster in the first clustering space and the second clustering space; removing the i-th type of derived data from the second clustering space, and obtaining the total number of users in the intersection of each cluster in the first clustering space and the second clustering space after removing the i-th type of derived data; where i is a positive integer greater than zero and less than or equal to the total number of types of derived data; based on the total number of users in the intersection of each cluster in the first clustering space and the second clustering space, and the total number of users in the intersection of each cluster in the first clustering space and the second clustering space after removing the i-th type of derived data, determining the weight value of the i-th type of derived data.
[0010] In a possible implementation of the first aspect, the verifying the reliability of the derived data output by the preset language model based on the first clustering result and the second clustering result includes: obtaining the similarity between the first clustering result and the second clustering result; based on the similarity between the first clustering result and the second clustering result, verifying the reliability of the derived data output by the preset language model.
[0011] In a possible implementation of the first aspect, the original data is user transaction data; the derived data includes at least one of sentiment data, trustworthiness data, suspiciousness data, importance data, and reliability data.
[0012] Second aspect, an embodiment of the present application provides a verification device for derivative data, including: an acquisition module, configured to acquire sample data; wherein, the sample data includes the original data and derivative data of multiple users; the derivative data is generated by a preset language model; the preset language model is used to analyze the input original data and generate the derivative data required by the user; a clustering module, configured to cluster the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain a first clustering result and a second clustering result; wherein, the first clustering space includes the original data of the multiple users; the second clustering space includes the derivative data of the multiple users; the first clustering result corresponds to the first clustering space; the second clustering result corresponds to the second clustering space; a verification module, configured to verify the reliability of the derivative data output by the preset language model based on the first clustering result and the second clustering result.
[0013] Third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the verification method for derivative data according to any one of the above first aspects when executing the computer program.
[0014] Fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program implements the verification method for derivative data according to any one of the above first aspects when executed by a processor.
[0015] Fifth aspect, an embodiment of the present application provides a computer program product, which causes a terminal device to execute the verification method for derivative data according to any one of the above first aspects when the computer program product runs on the terminal device.
[0016] It can be understood that the beneficial effects of the above second aspect to fifth aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here.
[0017] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: For the method for verifying derivative data provided by the embodiments of the present application, after obtaining the original data and derivative data of multiple users, clustering is performed on the data in the first clustering space (including the original data of multiple users) and the second clustering space (including the derivative data of multiple users) based on a preset clustering algorithm to obtain a first clustering result and a second clustering result. Then, based on the principle that "for the same batch of logical samples, the clustering results in different clustering spaces are similar", the reliability of the output of the preset language model is verified. It can be seen that through the above method, the reliability of the derivative data output by the preset language model can be effectively verified. At the same time, it is also convenient for users to perform subsequent analysis and processing on the derivative data. For example, when a user needs to perform risk control analysis, customer scoring, etc. based on transaction information, they can first use a conventional preset language model to obtain the expected derivative data, and after verifying the reliability of the derivative data through the above method, directly use the derivative data for risk control analysis, customer scoring, etc., without having to reconstruct a network model for analyzing transaction information; to a certain extent, it can also improve the processing efficiency of users. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 is a schematic structural diagram of a terminal device provided by an embodiment of the present application;
[0020] Figure 2 is a schematic flowchart of a method for verifying derivative data provided by an embodiment of the present application;
[0021] Figure 3 is a schematic flowchart of a method for verifying derivative data provided by another embodiment of the present application;
[0022] Figure 4 is a schematic flowchart of a method for verifying derivative data provided by still another embodiment of the present application;
[0023] Figure 5 is a schematic structural diagram of a device for verifying derivative data provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In the following description, specific details such as specific system architectures and technologies are presented for purposes of illustration and not limitation, so as to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present application.
[0025] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0026] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0027] As used in the specification of the present application and the appended claims, the term "if" may be interpreted, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0028] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are used only for differentiating descriptions and cannot be understood as indicating or implying relative importance.
[0029] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0030] With the development of technology, more and more chat tools based on language models are being used by people, such as the currently popular ChatGPT. ChatGPT is a natural language processing tool driven by artificial intelligence technology. It can generate answers based on the patterns and statistical laws seen during the pre-training stage and can also interact according to the context of the chat.
[0031] However, in various fields currently, there will be rich domain-specific information. How to use the above ChatGPT-like models to stably, reliably, and effectively understand this domain knowledge and embed it into the vector space is an important and novel research topic currently. In other words, there is currently a lack of a way to verify the reliability of the data derived from ChatGPT-like models.
[0032] Taking the user transaction information in the financial field as an example, when a user needs to perform risk control analysis on the transaction information, the user transaction data can be input into ChatGPT so that ChatGPT outputs the risk level or suspicious level of the transaction data. Then, the user can perform risk control analysis, but currently, it is impossible to know the reliability of the derived data (such as the above risk level data, suspicious level data) of the transaction data output by ChatGPT.
[0033] Therefore, the present application provides the following embodiments to solve the above technical problems.
[0034] Please refer to Figure 1 , the present application embodiment provides a module frame of a terminal device 1 for applying a verification method of derived data. The terminal device 1 includes: at least one processor 10 ( Figure 1 only one is shown in the figure), a memory 11, and a computer program 12 stored in the memory 11 and executable on at least one processor 10. When the processor 10 executes the computer program 12, it implements the steps in any of the above embodiments of the verification method of derived data.
[0035] The terminal device 1 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device 1 may include, but is not limited to, a processor 10 and a memory 11. Those skilled in the art can understand that Figure 1 merely an example of the terminal device, which does not constitute a limitation on the terminal device, and may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0036] The so-called processor 10 may be a Central Processing Unit (CPU), and the processor 10 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0037] In some embodiments, the memory 11 may be an internal storage unit of the terminal device 1, such as the hard disk or memory of the terminal device 1. In some other embodiments, the memory 11 may also be an external storage device of the terminal device 1, such as a plug-in hard disk equipped on the terminal device 1, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 11 may also include both the internal storage unit and the external storage device of the terminal device 1. The memory 11 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program, etc. The memory 11 may also be used to temporarily store data that has been output or will be output.
[0038] Please refer to Figure 2 , an embodiment of the present application provides a method for verifying derivative data. By way of example and not limitation, this method may be applied to the above-mentioned terminal device 1. This method may specifically include: step S201 - step S203.
[0039] Step S201: Obtain sample data.
[0040] Among them, the sample data includes the original data and derivative data of multiple users; the derivative data is generated by a preset language model; the preset language model is used to analyze the input original data and generate the derivative data required by the user.
[0041] The above-mentioned preset language module may be, but is not limited to, the chatGPT model, a chatGPT-like model (such as YouCHAT), and other neural network models constructed based on the Transformer model, which is not limited in this application.
[0042] It should be noted that the original data can be directly extracted from the information input by multiple users, while the derivative data can be data that cannot be directly extracted from the information input by users.
[0043] For example, if the data input by the user is the user's transaction information, the original data can be the account name, transaction amount, transaction time, transaction message, transfer description, remaining amount, transaction object, etc. in the user's transaction information; the above original data can be directly extracted from the information input by the user through text recognition; in other words, the original data can directly be the user's transaction data extracted from the user's transaction information. The derivative data can be at least one of emotional data, emotional data, trust degree data, suspicious degree data, importance degree data, and reliability degree data. The derivative data is generated by analyzing and processing the aforementioned user's transaction data.
[0044] It should be explained that: Emotional data is used to evaluate the emotional state shown by the user's transaction data, such as optimistic, pessimistic, neutral, etc. The acquisition of emotional data can be based on natural language processing technology to perform sentiment analysis on the text, extract keywords, or use a sentiment dictionary to judge the sentiment tendency.
[0045] Trust degree data is used to evaluate the degree of trust between the parties involved in the user's transaction data. It can be evaluated based on factors such as historical data and credit ratings, and the degree of trust is mapped to an appropriate numerical range.
[0046] Suspicious degree data is used to evaluate whether there are abnormal or suspicious situations in the user's transaction data. It can be evaluated based on factors such as transaction amount, transaction frequency, and transaction object, and the suspicious degree is mapped to an appropriate numerical range.
[0047] Importance degree data is used to evaluate the importance level or priority of the user's transaction data. It can be evaluated based on factors such as transaction type, amount size, and transaction object, and the importance degree is mapped to an appropriate numerical range.
[0048] Reliability degree data is used to evaluate the reliability degree of the user's transaction data, that is, the degree of trust in this record. It can be evaluated based on factors such as historical data and the background information of the parties involved, and the reliability degree is mapped to an appropriate numerical range.
[0049] Taking the above user's transaction information as an example, the following describes the complete process of obtaining sample data:
[0050] Here, first obtain the user's transaction information input by the user; then perform text processing on it to obtain the user's transaction data (i.e., the original data). Here, the user's transaction data can be processed into tabular data, that is, each row of data represents a transaction record in the user's transaction data. For the convenience of storage, each row of data can be embedded into a fixed field name space and value space.
[0051] It should be noted that when inputting the user's transaction information as described above, multiple historical transaction information of the user can be input to facilitate the subsequent generation of derivative data by the preset language model.
[0052] After obtaining the user's transaction data (i.e., the original data), it can be used as the input of the preset language model. It should be noted that the information input at this time also includes the derivative data expected by the user; for example, if the user expects to obtain sentiment data; then at this time, the user can input "the user's transaction data (i.e., the original data) and the instruction to inform the sentiment data". Taking the chatGPT model as an example, the user can input "Please analyze the sentiment data of the user's transaction data (i.e., the original data)" in the input box.
[0053] In this way, the derivative data corresponding to the original data can be obtained. Then, by organizing the original data and the derivative data, the sample data at this time can be obtained.
[0054] It should be noted that this method can be applied not only to the processing of transaction data in the financial field, but also to other specific fields. For example, this method can also be applied to the education field, in which case the above-mentioned original data can be student information; for example, the original data can specifically include student name, age, grades, hobbies, positions, etc. And the derivative data can specifically include the data on the adaptability of students in school, sentiment data (positive, negative, neutral), stability (such as grade stability), etc., which are not limited in this application.
[0055] Step S202: Cluster the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain a first clustering result and a second clustering result.
[0056] Among them, the first clustering space includes the original data of multiple users; the second clustering space includes the derivative data of multiple users; the first clustering result corresponds to the first clustering space; the second clustering result corresponds to the second clustering space.
[0057] The above-mentioned preset clustering algorithm can be, but is not limited to, the GMM (Gaussian Mixture Model) clustering algorithm, the K-Means clustering algorithm, etc.
[0058] Specifically, two clustering spaces can be initialized first, namely the first clustering space and the second clustering space. Then, the original data of multiple users is mapped to the first clustering space, and the derivative data of multiple users is mapped to the second clustering space. Then, the same preset clustering algorithm is used to cluster the data in the first clustering space and the second clustering space, and then the first clustering result and the second clustering result are obtained.
[0059] For example, the original data can include the account name, transaction amount, transaction time, transaction message, transfer description, remaining amount, and transaction object in the user transaction data. The derivative data can include sentiment data, sentiment data, trust degree data, suspicious degree data, importance degree data, and reliability data. The above data can all use the customer corresponding to the account name of the user transaction data as the clustering subject. Suppose the customers include customer A, customer B, customer C, and customer D. Then, the original data of customer A, the original data of customer B, the original data of customer C, and the original data of customer D are mapped to the first clustering space. At this time, each original data serves as the label data of the corresponding customer. The derivative data of customer A, the derivative data of customer B, the derivative data of customer C, and the derivative data of customer D are mapped to the second clustering space. At this time, each diffraction data serves as the label data of the corresponding customer. Then, the same preset clustering algorithm is used to cluster the data in the first clustering space and the second clustering space, and then the first clustering result and the second clustering result can be obtained. Among them, the first clustering result includes the clustering results of customer A, customer B, customer C, and customer D after analyzing the original data of the customers. The second clustering result includes the clustering results of customer A, customer B, customer C, and customer D after analyzing the derivative data of the customers.
[0060] Step S203: Based on the first clustering result and the second clustering result, verify the reliability of the derivative data output by the preset language model.
[0061] Finally, through the obtained first clustering result and second clustering result, the reliability of the derivative data output by the preset language model is further verified. It should be noted that in the embodiments of the present application, the principle of verifying the reliability of the derivative data output by the preset language model is: for the same batch of logical samples, the clustering results in different clustering spaces are similar.
[0062] In other words, when the clustering result obtained by analyzing the original data through the first clustering space is similar to the clustering result obtained by analyzing the derivative data through the second clustering space, it proves that the derivative data output by the preset language model is reliable. On the contrary, when the clustering result obtained by analyzing the original data through the first clustering space is not similar to the clustering result obtained by analyzing the derivative data through the second clustering space, it proves that the derivative data output by the preset language model is unreliable.
[0063] In summary, for the method for verifying derivative data provided in the embodiments of the present application, after obtaining the original data and derivative data of multiple users, the data in the first clustering space (including the original data of multiple users) and the second clustering space (including the derivative data of multiple users) are clustered based on a preset clustering algorithm to obtain a first clustering result and a second clustering result. Furthermore, based on the principle that "for the same batch of logical samples, the clustering results in different clustering spaces are similar", the reliability of the output of the preset language model is verified. It can be seen that through the above method, the reliability of the derivative data output by the preset language model can be effectively verified.
[0064] At the same time, it also facilitates the user to perform subsequent analysis and processing on the derivative data. When it is verified that the derivative data output by the preset language model is reliable, the derivative data output by the preset language model can be directly used for subsequent analysis and processing. For example, if the original data is user transaction data, after obtaining the derivative data through the output of the preset language model, the derivative data of each customer can be spliced with the original data, and then risk control analysis, customer scoring and other analysis and processing can be performed. After verifying the reliability of the derivative data through the above method, the derivative data can be directly used for risk control analysis, customer scoring, etc., without having to reconstruct a network model for analyzing transaction information; to a certain extent, it can also improve the processing efficiency of the user.
[0065] Optionally, in some embodiments, the step S202 of clustering the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain a first clustering result and a second clustering result may specifically include: performing clustering initialization on the first clustering space to obtain n clustering clusters, and generating n clustering clusters with the same categories in the second clustering space; n is a positive integer greater than 1; performing a clustering operation on the first clustering space to obtain the clustering result of the n clustering clusters in the first clustering space; randomly sampling each clustering cluster in the first clustering space to obtain a sampling result; using the sampling result as the initial clustering center of the n clustering clusters in the second clustering space, and performing a clustering operation on the second clustering space to obtain the clustering result of the n clustering clusters in the second clustering space; where the first clustering result includes the clustering result of the n clustering clusters in the first clustering space; the second clustering result includes the clustering result of the n clustering clusters in the second clustering space.
[0066] Specifically, two clustering spaces can be initialized first, namely the first clustering space and the second clustering space; then, the original data of multiple users is mapped to the first clustering space, and the derivative data of multiple users is mapped to the second clustering space. Then, perform clustering initialization on the first clustering space to obtain n clustering clusters; and generate n clustering clusters with the same categories in the second clustering space; where n is a positive integer greater than 1.
[0067] For example, the original data may include the account name, transaction amount, transaction time, transaction message, transfer description, remaining amount, and transaction object in the user transaction data. The derived data may include sentiment data, sentiment data, trustworthiness data, suspiciousness data, importance data, and reliability data. The above data may all use the customer corresponding to the account name of the user transaction data as the clustering subject. Suppose there are 100 customers. Then, map the original data of the 100 customers to the first clustering space. At this time, each piece of original data serves as the label data of the corresponding customer. Map the derived data of the 100 customers to the second clustering space. At this time, each piece of derived data serves as the label data of the corresponding customer. Then, perform clustering initialization on the first clustering space to obtain n clustering clusters. Here, the n clustering clusters obtained by clustering initialization can be set according to actual needs. For example, they can be classified according to customer levels to generate three clustering clusters: clustering cluster A, clustering cluster B, and clustering cluster C; clustering cluster A represents the set of type A customers, clustering cluster B represents the set of type B customers, and clustering cluster C represents the set of type C customers. Here, it can be that the higher the level, the better the customer and the higher the creditworthiness. Then, also generate three clustering clusters with the same categories in the second clustering space; namely, clustering cluster A, clustering cluster B, and clustering cluster C.
[0068] Next, use a preset clustering algorithm to perform a clustering operation on the first clustering space to obtain the clustering results of the n clustering clusters in the first clustering space. For example, when the preset clustering algorithm is the GMM clustering algorithm, this step can be to perform a single-step EM clustering operation on the first clustering space (the EM clustering operation is a clustering operation in the GMM clustering algorithm), and then obtain the clustering results of the n clustering clusters in the first clustering space. For example, the above 100 customers can be respectively divided into clustering cluster A, clustering cluster B, and clustering cluster C; for example, the clustering results of the first clustering space include: 30 customers corresponding to clustering cluster A, 60 customers corresponding to clustering cluster B, and 10 customers corresponding to clustering cluster C.
[0069] Next, perform random sampling on each clustering cluster in the first clustering space to obtain the sampling results. Here, random sampling means separately extracting a certain number of objects from each clustering cluster. It can be randomly extracted according to a proportion or a fixed number. For example, extract 5% of the objects in each clustering cluster, or fix 8 objects in each clustering cluster. It can also be randomly extracted with any value. In this regard, this application does not make a limitation. According to the previous example, it can be assumed that the sampling results at this time include extracting 6 customers from clustering cluster A, 10 customers from clustering cluster B, and 2 customers from clustering cluster C.
[0070] Next, use the sampling result as the initial cluster centers of the n clusters in the second clustering space, and perform a clustering operation on the second clustering space to obtain the clustering results of the n clusters in the second clustering space. For example, at this time, place the 6 customers extracted from cluster A in the first clustering space into cluster A in the second clustering space, place the 10 customers extracted from cluster B in the first clustering space into cluster B in the second clustering space, and place the 2 customers extracted from cluster C in the first clustering space into cluster C in the second clustering space, so as to endow the second clustering space with certain initial cluster centers, and use the same preset clustering algorithm to perform a clustering operation on the second clustering space. For example, when the preset clustering algorithm is the GMM clustering algorithm, this step can be to perform a single-step EM clustering operation on the second clustering space (the EM clustering operation is a clustering operation in the GMM clustering algorithm), and then obtain the clustering results of the n clusters in the second clustering space. By endowing the second clustering space with certain initial cluster centers, it can guide the second clustering space to cluster in the direction of the first clustering space and obtain the clustering results of the n clusters in the second clustering space. In other words, it is convenient to subsequently determine whether the second clustering space clusters in the clustering direction of the first clustering space, and it is also convenient to subsequently improve the accuracy of verifying whether the derivative data output by the preset language model is reliable.
[0071] In the above manner, the first clustering result and the second clustering result can be obtained; among them, the first clustering result includes the clustering results of the n clusters in the first clustering space; the second clustering result includes the clustering results of the n clusters in the second clustering space.
[0072] Please refer to Figure 3 , in some embodiments, the above step S202 clusters the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain the first clustering result and the second clustering result, and may specifically include: steps S301 - S308.
[0073] Step S301: Perform clustering initialization on the first clustering space to obtain n clusters, and generate n clusters with the same categories in the second clustering space.
[0074] Step S302: Perform a clustering operation on the first clustering space to obtain the clustering results of the n clusters in the first clustering space.
[0075] Step S303: Randomly sample each cluster in the first clustering space to obtain the sampling result of the first clustering space.
[0076] Step S304: Use the sampling result of the first clustering space as the initial cluster centers of the n clusters in the second clustering space, and perform a clustering operation on the second clustering space to obtain the clustering results of the n clusters in the second clustering space.
[0077] The specific processes of the above-mentioned steps S301 - S304 can refer to the descriptions in the foregoing embodiments. For the same parts, reference can be made to each other, and no detailed description will be given here.
[0078] Step S305: Randomly sample each cluster in the second clustering space to obtain the sampling result of the second clustering space.
[0079] Random sampling means separately extracting a certain number of objects from each cluster. It can be randomly extracting according to a proportion or extracting a fixed number. For example, extracting 5% of the objects in each cluster, or fixedly extracting 8 objects in each cluster. It can also be randomly extracting any value. In this regard, the present application makes no limitation.
[0080] Step S306: Use the sampling result of the second clustering space as the initial clustering centers of the n clusters in the first clustering space, and perform a clustering operation on the first clustering space to obtain the clustering result of the n clusters in the first clustering space.
[0081] It should be noted that steps S305 - S306 and steps S303 - S304 are corresponding reverse operations; steps S303 - S304 are to randomly sample in the first clustering space, use the sampling result of the first clustering space as the initial clustering center of the second clustering space, and perform a clustering operation on the second clustering space. While steps S305 - S306 are to randomly sample in the second clustering space, use the sampling result of the second clustering space as the initial clustering center of the first clustering space, and perform a clustering operation on the first clustering space again. Therefore, the specific details of the corresponding steps can be referred to each other, and no detailed description will be given here.
[0082] Step S307: Determine whether the clustering results of the n clusters in the first clustering space and the clustering results of the n clusters in the second clustering space converge.
[0083] If it has converged, execute step S308; if it has not converged, repeat the above steps S303 - S306 until the clustering results of the n clusters in the first clustering space and the clustering results of the n clusters in the second clustering space converge.
[0084] It should be noted that for the process of repeated iteration, for the sampling results, sampling with replacement or sampling without replacement can be used for continued sampling. The present application makes no limitation.
[0085] Here, convergence can indicate that the clustering results of the n clusters in the first clustering space and the clustering results of the n clusters in the second clustering space no longer change.
[0086] Step S308: Output the first clustering result and the second clustering result.
[0087] Among them, the first clustering result includes the clustering results of the n clustering clusters in the converged first clustering space; the second clustering result includes the clustering results of the n clustering clusters in the converged second clustering space.
[0088] In summary, the above embodiments provide a clustering algorithm with mutual verification in two spaces, that is, two vector spaces mutually verify, evolving from the EM algorithm in one vector space to the mutual verification of K Gaussian mixture models in two vector spaces. And through the above method, the accuracy of verifying whether the derivative data output by the preset language model is reliable can be further improved. In other words, through the clustering algorithm with mutual verification in two spaces, the similarity of the clustering results of the two clustering spaces can be accurately judged, reducing the possibility of misjudgment.
[0089] Optionally, in some embodiments, after verifying the reliability of the derivative data output by the preset language model, the method may further specifically include: determining the weight values of each derivative data; aggregating the original data of multiple users, the derivative data of multiple users, and the weight values of the derivative data of multiple users to generate target data.
[0090] It should be noted that considering the unknown unreliability inside the preset language model, when using each derivative data subsequently, the weight value of each derivative data needs to be combined. Then, aggregate the original data of multiple users, the derivative data of multiple users, and the weight values of the derivative data of multiple users to generate target data. Finally, use the target data for subsequent processing and analysis.
[0091] For example, if the original data is user transaction data, after obtaining the derivative data through the output of the preset language model, the derivative data of each customer (i.e., the above-mentioned user) can be multiplied by the weight value, then concatenated with the original data to obtain the target data, and finally use the target data for risk control analysis, customer scoring and other analysis processes.
[0092] Here, the derivative data can also be embedded into a fixed field name space and a derivative data value space. Then it is concatenated with the fixed field name space and value space of the original data.
[0093] Optionally, in some embodiments, for the verification process, it may further include checking whether the output is a dataframe in a fixed field name space and verifying whether the data format of each field name is consistent with the setting.
[0094] Please refer to Figure 4 , optionally, the weight values of each derivative data can be determined through the following steps, including: Step S401 - Step S403.
[0095] Step S401: Obtain the total number of users in the intersection of each cluster in the first clustering space and the second clustering space.
[0096] The total number of users in the above intersection can be calculated by the following formula: where K represents the total number of clusters; U Aj represents the number of users in the j-th cluster in the first clustering space (A space); U Bj represents the number of users in the j-th cluster in the second clustering space (B space); the above formula represents the total number of users in the intersection of each cluster in the first clustering space and the second clustering space.
[0097] Step S402: Remove the i-th type of derived data from the second clustering space, and obtain the total number of users in the intersection of each cluster in the first clustering space and the second clustering space after removing the i-th type of derived data.
[0098] where i is a positive integer greater than zero and less than or equal to the total number of types of derived data.
[0099] It should be noted that i starts from 1 and increases sequentially to remove each type of derived data in turn, so as to determine the weight value of each derived feature.
[0100] Then, repeat the calculation in step S402 N times, where N is the total number of types of derived data.
[0101] Step S403: Based on the total number of users in the intersection of each cluster in the first clustering space and the second clustering space, and the total number of users in the intersection of each cluster in the first clustering space and the second clustering space after removing the i-th type of derived data, determine the weight value of the i-th type of derived data.
[0102] The specific formula for the weight value includes: where w i represents the weight value of the i-th type of derived data. where Δ|x i | represents the difference between the total number of users in the intersection of each cluster of the two spaces after removing the i-th type of derived data from the second clustering space and the reference number of people. The reference number of people is the total number of users in the intersection of each cluster in the first clustering space and the second clustering space in step S401. x i represents the influence of the i-th type of derived data on the overall clustering effect.
[0103] In summary, the calculation of the above weight value considers the influence on the overall clustering effect after each type of derived data is removed. The greater the influence, the higher the weight value; on the contrary, the smaller the influence, the lower the weight value.
[0104] Of course, in other embodiments, the weight value can also be set according to the user's own needs, and this application does not make any limitations in this regard.
[0105] Optionally, in some embodiments, the above step S203 verifies the reliability of the derivative data output by the preset language model based on the first clustering result and the second clustering result, including: obtaining the similarity between the first clustering result and the second clustering result; and verifying the reliability of the derivative data output by the preset language model based on the similarity between the first clustering result and the second clustering result.
[0106] Here, it can be that when the similarity between the first clustering result and the second clustering result is greater than a preset value, it indicates that the derivative data output by the preset language model is reliable; for example, the above preset can be 80%, 90%, 95%, etc., and this application does not make any limitations.
[0107] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0108] Corresponding to the derivative data verification method described in the above embodiments, Figure 5 The structural block diagram of the derivative data verification device 500 provided by the embodiments of this application is shown. For the sake of convenience of description, only the parts related to the embodiments of this application are shown.
[0109] Please refer to Figure 5 , the device includes: an acquisition module 501, configured to acquire sample data; wherein, the sample data includes the original data and derivative data of multiple users; the derivative data is generated by a preset language model; the preset language model is used to analyze the input original data and generate the derivative data required by the user; a clustering module 502, configured to cluster the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain a first clustering result and a second clustering result; wherein, the first clustering space includes the original data of the multiple users; the second clustering space includes the derivative data of the multiple users; the first clustering result corresponds to the first clustering space; the second clustering result corresponds to the second clustering space; a verification module 503, configured to verify the reliability of the derivative data output by the preset language model based on the first clustering result and the second clustering result.
[0110] Optionally, in some embodiments, the clustering module 502 may further be specifically configured to perform clustering initialization on the first clustering space to obtain n clustering clusters, and generate n clustering clusters with the same categories in the second clustering space; n is a positive integer greater than 1; perform a clustering operation on the first clustering space to obtain the clustering results of the n clustering clusters in the first clustering space; randomly sample each clustering cluster in the first clustering space to obtain a sampling result; use the sampling result as the initial clustering centers of the n clustering clusters in the second clustering space, and perform a clustering operation on the second clustering space to obtain the clustering results of the n clustering clusters in the second clustering space; wherein, the first clustering result includes the clustering results of the n clustering clusters in the first clustering space; the second clustering result includes the clustering results of the n clustering clusters in the second clustering space.
[0111] Optionally, in some embodiments, the clustering module 502 may further be specifically configured to: S1: perform clustering initialization on the first clustering space to obtain n clustering clusters, and generate n clustering clusters with the same categories in the second clustering space; n is a positive integer greater than 1; S2: perform a clustering operation on the first clustering space to obtain the clustering results of the n clustering clusters in the first clustering space; S3: randomly sample each clustering cluster in the first clustering space to obtain the sampling result of the first clustering space; S4: use the sampling result of the first clustering space as the initial clustering centers of the n clustering clusters in the second clustering space, and perform a clustering operation on the second clustering space to obtain the clustering results of the n clustering clusters in the second clustering space; S5: randomly sample each clustering cluster in the second clustering space to obtain the sampling result of the second clustering space; S6: use the sampling result of the second clustering space as the initial clustering centers of the n clustering clusters in the first clustering space, and perform a clustering operation on the first clustering space to obtain the clustering results of the n clustering clusters in the first clustering space; S7: repeat the above steps S3 - S6 until the clustering results of the n clustering clusters in the first clustering space and the clustering results of the n clustering clusters in the second clustering space converge; S8: output the first clustering result and the second clustering result; wherein, the first clustering result includes the clustering results of the n clustering clusters in the first clustering space after convergence; the second clustering result includes the clustering results of the n clustering clusters in the second clustering space after convergence.
[0112] Optionally, in some embodiments, the apparatus further includes: a generation module. The generation module is configured to determine the weight values of the respective derived data after verifying that the reliability of the derived data output by the preset language model is passed; summarize the original data of the multiple users, the derived data of the multiple users, and the weight values of the derived data of the multiple users to generate target data.
[0113] Optionally, in some embodiments, the generating module is further specifically configured to obtain the total number of users in the intersection of each cluster in the first clustering space and the second clustering space; remove the i-th type of derivative data from the second clustering space, and obtain the total number of users in the intersection of each cluster in the first clustering space and the second clustering space after removing the i-th type of derivative data; where i is a positive integer greater than zero and less than or equal to the total number of derivative data; based on the total number of users in the intersection of each cluster in the first clustering space and the second clustering space, and the total number of users in the intersection of each cluster in the first clustering space and the second clustering space after removing the i-th type of derivative data, determine the weight value of the i-th type of derivative data.
[0114] Optionally, in some embodiments, the verification module 503 is further specifically configured to obtain the similarity between the first clustering result and the second clustering result; based on the similarity between the first clustering result and the second clustering result, verify the reliability of the derivative data output by the preset language model.
[0115] It should be noted that, for the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0116] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used for illustration. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, and details are not described herein again.
[0117] The embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0118] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, it enables the mobile terminal to execute and implement the steps in the above-mentioned method embodiments.
[0119] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps in the above-mentioned method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0120] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0121] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0122] In the embodiments provided in the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical or other forms.
[0123] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0124] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for verifying derivative data, characterized in that, Including: Obtain sample data; wherein, the sample data includes the original data and derivative data of multiple users; the derivative data is generated by a preset language model; the preset language model is used to analyze the input original data and generate the derivative data required by the user, and the original data is user transaction data, which is data directly extracted from the information input by the user through text recognition; the derivative data includes at least one of sentiment data, trust degree data, suspicious degree data, importance data, and reliability data; Cluster the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain a first clustering result and a second clustering result; wherein, the first clustering space includes the original data of the multiple users; the second clustering space includes the derivative data of the multiple users; the first clustering result corresponds to the first clustering space; the second clustering result corresponds to the second clustering space; both the original data and the derivative data use the customer corresponding to the account name of the user transaction data as the clustering subject; Verify the reliability of the derivative data output by the preset language model based on the first clustering result and the second clustering result.
2. The method for verifying derivative data according to claim 1, wherein, The clustering of the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain a first clustering result and a second clustering result includes: Perform clustering initialization on the first clustering space to obtain n clustering clusters, and generate n clustering clusters with the same category in the second clustering space; n is a positive integer greater than 1; Perform a clustering operation on the first clustering space to obtain the clustering result of the n clustering clusters in the first clustering space; Randomly sample each clustering cluster in the first clustering space to obtain a sampling result; Use the sampling result as the initial clustering center of the n clustering clusters in the second clustering space, and perform a clustering operation on the second clustering space to obtain the clustering result of the n clustering clusters in the second clustering space; Wherein, the first clustering result includes the clustering result of the n clustering clusters in the first clustering space; the second clustering result includes the clustering result of the n clustering clusters in the second clustering space.
3. The method for verifying derivative data according to claim 1, characterized in that, The clustering of the data in the first clustering space and the second clustering space based on a preset clustering algorithm to obtain a first clustering result and a second clustering result includes: S1: Perform clustering initialization on the first clustering space to obtain n clustering clusters, and generate n clustering clusters with the same category in the second clustering space; n is a positive integer greater than 1; S2: Perform a clustering operation on the first clustering space to obtain the clustering result of the n clustering clusters in the first clustering space; S3: Randomly sample each clustering cluster in the first clustering space to obtain the sampling result of the first clustering space; S4: Use the sampling result of the first clustering space as the initial clustering center of the n clustering clusters in the second clustering space, and perform a clustering operation on the second clustering space to obtain the clustering result of the n clustering clusters in the second clustering space; S5: Randomly sample each cluster in the second clustering space to obtain the sampling result of the second clustering space; S6: Use the sampling result of the second clustering space as the initial cluster centers of the n clusters in the first clustering space, and perform a clustering operation on the first clustering space to obtain the clustering result of the n clusters in the first clustering space; S7: Repeat the above steps S3 - S6 until the clustering results of the n clusters in the first clustering space and the clustering results of the n clusters in the second clustering space converge; S8: Output the first clustering result and the second clustering result; wherein, the first clustering result includes the clustering results of the n clusters in the first clustering space after convergence; the second clustering result includes the clustering results of the n clusters in the second clustering space after convergence.
4. The verification method of the derivative data according to any one of claims 2-3, characterized in that, After verifying that the reliability of the derivative data output by the preset language model passes, the method further includes: Determine the weight values of the respective derivative data; Summarize the original data of the multiple users, the derivative data of the multiple users, and the weight values of the derivative data of the multiple users to generate target data.
5. The method for verifying derivative data according to claim 4, wherein The determining the weight values of the respective derivative data includes: Obtain the total number of user counts in the intersection of each cluster in the first clustering space and the second clustering space; Remove the i-th type of derivative data from the second clustering space, and obtain the total number of user counts in the intersection of each cluster in the first clustering space and the second clustering space after removing the i-th type of derivative data; where i is a positive integer greater than zero and less than or equal to the total number of derivative data; Based on the total number of user counts in the intersection of each cluster in the first clustering space and the second clustering space, and the total number of user counts in the intersection of each cluster in the first clustering space and the second clustering space after removing the i-th type of derivative data, determine the weight value of the i-th type of derivative data.
6. The method for verifying derivative data according to claim 1, wherein, The verifying the reliability of the derivative data output by the preset language model based on the first clustering result and the second clustering result includes: Obtain the similarity between the first clustering result and the second clustering result; Based on the similarity between the first clustering result and the second clustering result, verify the reliability of the derivative data output by the preset language model.
7. A verification device for derivative data, characterized in that, Includes: An acquisition module for acquiring sample data; wherein, the sample data includes the original data and derivative data of multiple users; the derivative data is generated by a preset language model; the preset language model is used to analyze the input original data to generate the derivative data required by the user, the original data is user transaction data, which is directly extracted from the information input by the user through text recognition; the derivative data includes at least one of sentiment data, trust degree data, suspiciousness data, importance data, and reliability data; A clustering module, configured to cluster data in a first clustering space and a second clustering space based on a preset clustering algorithm, to obtain a first clustering result and a second clustering result; wherein, the first clustering space includes the original data of the multiple users; the second clustering space includes the derived data of the multiple users; the first clustering result corresponds to the first clustering space; the second clustering result corresponds to the second clustering space; both the original data and the derived data use the customers corresponding to the account names of the user transaction data as the clustering entities; A verification module, configured to verify the reliability of the derived data output by the preset language model based on the first clustering result and the second clustering result.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Wealth trend prediction method and device, equipment and storage medium
CN111242356A
Power consumer value analysis method and system based on payment behavior
CN112990721A