Recommendation platform data verification method and related equipment
By performing data consistency checksum attribution analysis on the offline data link of the recommended platform, the problem of low data verification efficiency and accuracy in the existing technology is solved, and efficient data verification problem investigation is achieved.
Patent Information
- Application Number
- CN202510252605.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-06
AI Technical Summary
The data verification of existing recommended platforms relies on manual operations, with low efficiency and accuracy, and due to the differences in different data types and formats, the complexity increases, resulting in difficulty in problem location and root cause traceability.
The data of each node on the offline data link of the recommendation platform is read through a storage path based on the baseline data and test data, data consistency verification is performed, the difference data is determined, and attribution analysis is performed to determine the node generating the difference data.
It realizes efficient data verification of offline data links on the recommendation platform, improves problem detection efficiency, reduces the workload and probability of errors of manual intervention, and provides higher quality differential data analysis results.
Smart Images

Figure CN120104500A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a recommendation platform data verification method and related equipment. Background Art
[0002] The offline data link correctness verification of the recommendation platform (referred to as data verification) is a key link to ensure the stability of the recommendation platform system and data consistency, which is especially important in scenarios such as environment migration, component switching, and data cloud migration.
[0003] At present, the data verification of the recommendation platform mainly relies on manual operation, which is a lengthy process and has a large amount of repetitive work, resulting in low efficiency and accuracy of data verification. In addition, the data storage and processing of the recommendation platform usually covers a variety of data types and formats, such as batch data, streaming data, and key-value (KV) data. Among them, the data verification methods and technical implementations of different types and formats are quite different, which further aggravates the complexity of data verification of the recommendation platform. Furthermore, the offline data link of the recommendation platform usually involves multiple production links, and these production links have the characteristics of front-to-back series connection. The difference of upstream data will be transmitted to the downstream, making it difficult to determine whether the abnormal data downstream is its own problem or the influence of upstream transmission. This difference transmission mechanism makes problem location and root cause tracing more complicated, further increasing the difficulty of problem troubleshooting. Based on the above reasons, how to implement data verification of the recommendation platform has become one of the key issues that need to be solved to maintain the normal operation of the recommendation platform. Summary of the invention
[0004] In view of this, an embodiment of the present disclosure provides a recommendation platform data verification method, which can efficiently and accurately verify the data of the recommendation platform offline data link, thereby realizing efficient problem troubleshooting of the recommendation platform offline data link.
[0005] The recommendation platform data verification method described in the embodiment of the present disclosure includes: reading the baseline data corresponding to each node on the recommendation platform offline data link based on the storage path of the baseline data; reading the test data corresponding to each node based on the storage path of the test data; performing data consistency verification on the test data corresponding to each node based on the baseline data corresponding to each node, and respectively determining the difference data corresponding to each node reflecting the difference between the test data and the baseline data; wherein the difference data includes at least one difference data item; performing attribution analysis on the difference data corresponding to each node, and determining the analysis results of the difference data corresponding to each node; wherein the analysis results include: the node on the recommendation platform offline data link that generates each difference data item in the difference data; and displaying the analysis results of the difference data corresponding to each node.
[0006] In an embodiment of the present disclosure, the baseline data corresponding to each node is a baseline data table; the test data corresponding to each node is a test data table; the data consistency check of the test data corresponding to each node based on the baseline data corresponding to each node includes: for a target node on the offline data link of the recommendation platform, a row-level merge is performed on the test data table corresponding to the target node and the benchmark data table corresponding to the target node through a primary key-based merge operation; data comparison is performed on the fields corresponding to the same primary key to obtain row-level difference data items between the test data table corresponding to the target node and the benchmark data table corresponding to the target node; and the row-level difference data items are combined into difference data corresponding to the target node.
[0007] In an embodiment of the present disclosure, the baseline data corresponding to each node is batch data, streaming data or offline key-value data in a non-data table format; the test data corresponding to each node is batch data, streaming data or offline key-value data in a non-data table format; the data consistency check of the test data corresponding to each node based on the baseline data corresponding to each node includes: for a target node on the offline data link of the recommendation platform, converting the baseline data corresponding to the target node into a baseline data table corresponding to the target node, and converting the test data corresponding to the target node into a test data table corresponding to the target node; performing a row-level merge on the test data table corresponding to the target node and the benchmark data table corresponding to the target node through a primary key-based merge operation; performing data comparison on fields corresponding to the same primary key to obtain row-level difference data items between the test data table corresponding to the target node and the benchmark data table corresponding to the target node; and combining the row-level difference data items into difference data corresponding to the target node.
[0008] In an embodiment of the present disclosure, the baseline data corresponding to each node is online key-value data; the test data corresponding to each node is online key-value data; the data consistency check of the test data corresponding to each node based on the baseline data corresponding to each node includes: for a target node on the offline data link of the recommendation platform, sampling the key set of the baseline data corresponding to the target node to obtain the sampling key set corresponding to the target node; and performing the following verification operations on each sampling key in the sampling key set corresponding to the target node: performing a point query operation on the test data corresponding to the target node, and obtaining the sampling test value corresponding to the sampling key from the test data corresponding to the target node; performing a point query operation on the baseline data corresponding to the target node, and obtaining the sampling baseline value corresponding to the sampling key from the baseline data corresponding to the target node; comparing the sampling test value with the sampling baseline value; in response to determining that the sampling test value and the sampling baseline value are consistent, the data consistency check is successful; and in response to determining that the sampling test value and the sampling baseline value are inconsistent, recording the sampling key, the sampling test value and the sampling baseline value as a difference data item in the difference data corresponding to the target node.
[0009] In an embodiment of the present disclosure, the attribution analysis of the difference data corresponding to the various nodes includes: respectively extracting the primary key sets of the difference data corresponding to the various nodes; taking the various nodes as target nodes in turn according to the order of the nodes on the offline data link of the recommendation platform, and respectively performing the following operations: comparing the first primary key set of the difference data corresponding to the target node with the second primary key set of the difference data corresponding to the previous level node of the target node, determining the target primary key that exists in the first primary key set but does not exist in the second primary key set; and attributing the difference data item corresponding to the target primary key in the difference data corresponding to the target node to the target node.
[0010] In an embodiment of the present disclosure, the attribution analysis of the difference data corresponding to each node includes: taking each node as a target node in turn according to the order of the nodes on the offline data link of the recommendation platform, and performing the following operations respectively: generating a prompt of a large language model based on the difference data corresponding to the target node, the difference data of the node before the target node, and the data production logic of the target node; and inputting the prompt into the large language model, and the large language model outputs an attribution result of whether each difference data item in the difference data corresponding to the target node is attributed to the target node.
[0011] In an embodiment of the present disclosure, the above method further includes: performing statistical analysis on the difference data corresponding to each node to obtain the difference rate between the test data and the baseline data corresponding to each node; and adding the difference rate to the analysis result of the difference data corresponding to its corresponding node.
[0012] Corresponding to the above-mentioned recommendation platform data verification method, the embodiment of the present disclosure further discloses a recommendation platform data verification system, including:
[0013] A data storage module, used to store baseline data corresponding to each node on the offline data link of the recommendation platform, test data corresponding to each node, and difference data corresponding to each node reflecting the difference between the test data and the baseline data;
[0014] A configuration module, used to receive configuration information of a data verification task; wherein the configuration information includes: a storage path of baseline data, a storage path of test data, and a storage path of difference data;
[0015] a data verification module, configured to read the baseline data corresponding to each node from the data storage module based on the storage path of the baseline data, read the test data corresponding to each node from the data storage module based on the storage path of the test data, perform data consistency verification on the test data corresponding to each node based on the baseline data corresponding to each node, respectively determine the difference data corresponding to each node, and store the difference data corresponding to each node in the data storage module; wherein the difference data includes at least one difference data item;
[0016] a difference data analysis module, configured to read the difference data corresponding to each node from the data storage module, perform attribution analysis on the difference data corresponding to each node, and determine the analysis results of the difference data corresponding to each node; wherein the analysis results include: a node on the offline data link of the recommendation platform that generates each difference data item in the difference data; and
[0017] The analysis result display module is used to display the analysis results of the difference data corresponding to each node.
[0018] The above-mentioned recommendation platform data verification system may further include:
[0019] The scheduling module is used to pass the configuration information to the data verification module, schedule the data verification module to perform the data consistency verification task, and after the data verification module is executed, pass the configuration information to the difference data analysis module, and schedule the difference data analysis module to perform the difference data analysis task.
[0020] In addition, an embodiment of the present disclosure further provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned recommended platform data verification method when executing the program.
[0021] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the above-mentioned recommendation platform data verification method.
[0022] An embodiment of the present disclosure further provides a computer program product, including computer program instructions, which, when executed on a computer, enable the computer to execute the above-mentioned recommendation platform data verification method.
[0023] It can be seen from this that the recommendation platform data verification method and related equipment provided by some embodiments of the present disclosure can quickly verify the data of the recommendation platform offline data link through automated data difference troubleshooting, discover and locate data differences, significantly improve the efficiency of data verification, and reduce the workload and error probability of manual intervention. Furthermore, the above-mentioned recommendation platform data verification system can effectively eliminate the impact of data differences transmitted from the upstream links through the upstream and downstream data difference attribution mechanism, avoid the limitation of traditional tools that cannot accurately locate the root cause of the difference data, and provide higher quality difference data analysis results. In addition, the above-mentioned recommendation platform data verification system can also support multiple data types and multiple data sources, including batch data, streaming data, key-value data, etc., and can flexibly respond to the complex data flow environment of the recommendation platform and meet the needs of the recommendation platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0025] Figure 1 The internal structure of the recommendation platform data verification system described in some embodiments of the present disclosure is shown.
[0026] Figure 2 A timing diagram showing the collaborative completion of a recommendation platform data verification process by various modules in the recommendation platform data verification system described in some embodiments of the present disclosure.
[0027] Figure 3The relationship between the baseline data, test data, and difference data described in some embodiments of the present disclosure and the nodes on the offline data link of the recommendation platform is displayed.
[0028] Figure 4 An example of difference data corresponding to multiple nodes on an offline data link of a recommendation platform according to some embodiments of the present disclosure is shown.
[0029] Figure 5 An example of performing difference data attribution based on a large language model by the difference data attribution unit described in some embodiments of the present disclosure is shown.
[0030] Figure 6 The implementation process of the recommendation platform data verification method described in the embodiment of the present disclosure is shown.
[0031] Figure 7 A more specific schematic diagram of the hardware structure of an electronic device described in some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0032] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.
[0033] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0034] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0035] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information. Thus, the user can independently choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.
[0036] As an optional but non-limiting implementation, in response to receiving an active request from the user, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0037] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0038] As mentioned above, data verification of the recommendation platform is a key link to ensure the stability of the recommendation platform system and data consistency. However, there is currently a lack of efficient and accurate data verification solutions for the recommendation platform.
[0039] In order to solve the above problems, the embodiment of the present disclosure provides a recommendation platform data verification system, whose internal structure can be as follows: Figure 1 As shown, it can mainly include the following multiple modules: a data storage module 110 , a configuration module 120 , a data verification module 130 , a difference data analysis module 140 and an analysis result display module 150 .
[0040] In an embodiment of the present disclosure, the above-mentioned data storage module 110 can be used to store test data corresponding to each node on the offline data link of the recommendation platform, baseline data corresponding to each node on the offline data link of the recommendation platform, and data verification results corresponding to each node on the offline data link of the recommendation platform.
[0041] In the embodiments of the present disclosure, the test data is the data to be tested in the data verification; the baseline data is the reference data in the data verification. One of the goals of the data verification described in the embodiments of the present disclosure is to verify whether the data to be tested is consistent with the baseline data, or to find data in the test data that is inconsistent with the baseline data.
[0042] As mentioned above, the offline data link of the recommendation platform usually involves multiple production links, and these production links have the characteristics of front-to-back series connection. For ease of description, in the embodiment of the present disclosure, the multiple production links involved in the offline data link of the recommendation platform are abstracted as multiple nodes on the offline data link of the recommendation platform, and there is a series upstream and downstream relationship between these nodes. Therefore, in the embodiment of the present disclosure, the upstream node directly connected to a node can be referred to as its upper-level node, and the downstream node directly connected to it can be referred to as its lower-level node. In this way, another goal of the data verification described in the embodiment of the present disclosure is to attribute the data in the test data that is inconsistent with the baseline data, and determine which node on the offline data link of the recommendation platform generates the above-mentioned data inconsistency.
[0043] In addition, it can be understood that each node on the offline data link of the recommendation platform will perform certain processing on the data, so the data involved in each node will also be different, for example, including the data before and after processing, etc. Based on this, in the embodiments of the present disclosure, for the sake of clarity in description, the data obtained after processing by a node is referred to as the data corresponding to this node. Therefore, in the embodiments of the present disclosure, the test data, baseline data, and data verification results stored in the above-mentioned data storage module 110 all include the test data, baseline data, and data verification results corresponding to each node on the offline data link of the recommendation platform.
[0044] In addition, since the data storage and processing of the recommendation platform covers a variety of data types and data formats, in the embodiments of the present disclosure, the test data and baseline data also cover a variety of data types. For example, the test data and baseline data can be batch data, streaming data, or key-value structure data, etc. It can be understood that in a specific embodiment, the test data and baseline data should be data of the same type.
[0045] In addition, in an embodiment of the present disclosure, the above data verification result may include difference data reflecting the difference between the test data and the baseline data. Further, the above data verification result may also include: analysis results corresponding to the difference data obtained after analyzing the above difference data, such as attribution analysis results and / or statistical analysis results, etc.
[0046] In an embodiment of the present disclosure, the configuration module 120 may be used to receive configuration information of a data verification task.
[0047] In the embodiments of the present disclosure, the above configuration information may include: a storage path for baseline data, a storage path for test data, and a storage path for data verification results, etc. In addition, the above configuration information may also include: configuration information required for data consistency verification, such as a data sampling rate, etc.; and configuration information required for data analysis.
[0048] In an embodiment of the present disclosure, the above-mentioned data verification module 130 is used to read the baseline data corresponding to each node from the data storage module 110 based on the storage path of the baseline data, read the test data corresponding to each node from the data storage module 110 based on the storage path of the test data, perform data consistency verification on the test data corresponding to each node based on the baseline data corresponding to each node, respectively determine the difference data corresponding to each node, and store the difference data corresponding to each node in the data storage module 110.
[0049] Since both the test data and the baseline data usually include multiple data items, the difference data usually includes at least one difference data item. In some specific examples, each difference data item includes a test data item with a difference and a baseline data item corresponding thereto.
[0050] In the embodiment of the present disclosure, the difference data analysis module 140 is used to read the difference data corresponding to each node from the data storage module 110, perform attribution analysis on the difference data corresponding to each node, and determine the analysis result of the difference data corresponding to each node.
[0051] In addition, the difference data analysis module 140 may further submit analysis results corresponding to the difference data corresponding to each node to the data storage module 110 for storage.
[0052] In some specific examples, the analysis results of the difference data corresponding to the above-mentioned nodes may include: the nodes on the recommendation platform offline data link that generate each difference data item in the difference data. It can be seen that the above-mentioned analysis results can reflect the nodes that generate each difference data item or the production links that generate each difference data item, thereby completing the attribution of the difference data.
[0053] In the embodiment of the present disclosure, the analysis result display module 150 is used to display the analysis results of the difference data corresponding to each node. Specifically, the analysis result display module 150 can provide a data dashboard for the user, and display the analysis results corresponding to the difference data to the user in a visual manner.
[0054] In an embodiment of the present disclosure, the above-mentioned recommendation platform data verification system may further include: a scheduling module, used to schedule the above-mentioned data verification module 130 to perform data consistency verification based on the configuration information from the above-mentioned configuration module 120, and to schedule the above-mentioned difference data analysis module 140 to perform data analysis on the difference data based on the configuration information from the above-mentioned configuration module 120.
[0055] Specifically, after the user inputs the configuration information, the above-mentioned scheduling module can pass various types of configuration information to the data verification module 130, and schedule the data verification module 130 to perform data consistency verification tasks to verify different types of test data, and store the difference data obtained through verification in the data storage module 110. The above-mentioned scheduling module can also pass various types of configuration information to the difference data analysis module 140 after the task of the data verification module 130 is completed, and schedule the difference data analysis module 140 to perform the difference data analysis task to extract the difference data and perform statistical analysis and attribution analysis, and provide the analysis results to the analysis result display module 150 for displaying the final analysis results. In actual applications, the above-mentioned scheduling module can be implemented by applying existing thread scheduling tools.
[0056] Figure 2 The timing diagram of the recommendation platform data verification process in which each module in the recommendation platform data verification system according to the embodiment of the present disclosure cooperates to complete the recommendation platform data verification process is shown. Figure 2 As shown, the data verification process described in the embodiment of the present disclosure mainly includes the following steps.
[0057] In step 205 , the configuration module 120 receives configuration information of the data verification task from the user.
[0058] In step 210 , the configuration module 120 submits the data verification task including the configuration information to the scheduling module.
[0059] In step 215 , the scheduling module schedules the data verification module 130 to perform data consistency verification based on the configuration information.
[0060] In step 220, the data verification module 130 reads the baseline data corresponding to each node from the data storage module 110 based on the storage path of the baseline data in the configuration information, and reads the test data corresponding to each node from the data storage module 110 based on the storage path of the test data in the configuration information.
[0061] In step 225 , the data storage module 110 returns the baseline data corresponding to each node and the test data corresponding to each node to the data verification module 130 .
[0062] In step 230 , the data verification module 130 performs data consistency verification on the read test data based on the read baseline data, and determines the difference data corresponding to each node reflecting the difference between the test data and the baseline data.
[0063] In step 235 , the data verification module 130 stores the difference data corresponding to each node in the data storage module 110 .
[0064] In step 240, the data verification module 130 notifies the scheduling module that the data consistency verification is completed.
[0065] In step 245 , the scheduling module schedules the difference data analysis module 140 to perform data analysis on the difference data corresponding to each node based on the configuration information.
[0066] In step 250 , the difference data analysis module 140 reads the difference data corresponding to each node from the data storage module 110 .
[0067] In step 255 , the data storage module 110 returns the difference data corresponding to each node to the difference data analysis module 140 .
[0068] In step 260 , the difference data analysis module 140 performs attribution analysis on the difference data corresponding to each node to obtain analysis results of the difference data corresponding to each node.
[0069] In step 265 , the difference data analysis module 140 submits the analysis results of the difference data corresponding to each node to the analysis result display module 150 .
[0070] In step 270 , the analysis result display module 150 displays the analysis results of the difference data corresponding to each node to the user.
[0071] In some embodiments of the present disclosure, before executing the above step 270, the analysis result display module 150 may also cache the analysis results of the difference data corresponding to each of the above nodes, and then execute the above step 270 after receiving an analysis result display request from the user.
[0072] The specific implementation of the data consistency verification by the data verification module 130 and the data analysis by the difference data analysis module 140 in the above-mentioned recommendation platform data verification system will be described in detail below with reference to specific examples.
[0073] As mentioned above, in the embodiment of the present disclosure, the test data and baseline data can be one of multiple types of data such as batch data, streaming data or key-value structure data. Therefore, in the embodiment of the present disclosure, the data verification module 130 will use different processing methods to verify different types of data.
[0074] Specifically, in the embodiments of the present disclosure, the data verification module 130 may specifically include: a batch data processing unit, a streaming data processing unit, and a key-value data processing unit. The batch data processing unit is mainly used to perform data consistency verification on test data of batch data types based on baseline data; the streaming data processing unit is mainly used to perform data consistency verification on test data of streaming data types based on baseline data; and the key-value data processing unit is mainly used to perform data consistency verification on test data of key-value data types based on baseline data.
[0075] Among them, batch data also includes batch data in data table format, such as Hive table data, etc., and batch data in non-data table format, such as Feature Store sample data, etc.
[0076] For batch data in the form of data tables, that is, the above-mentioned test data and baseline data are both in the form of data tables, which are respectively referred to as test data tables and baseline data tables. The above-mentioned batch data processing unit can target a certain target node on the offline data link of the recommendation platform, and first perform a row-level merge on the test data table corresponding to the read target node and the benchmark data table corresponding to the target node through a primary key-based merge (Join) operation; then perform data comparison on the specific fields corresponding to the same primary key, so as to obtain row-level difference data items between the test data table corresponding to the target node and the benchmark data table corresponding to the target node; and finally combine the row-level difference data items into difference data corresponding to the target node.
[0077] Here, the row-level difference data items may include three situations: missing rows, redundant rows, and inconsistent rows. The missing rows refer to row data in the baseline data table that exists in the baseline data table but does not exist in the test data table; the redundant rows refer to row data in the test data table that exists in the test data table but does not exist in the baseline data table; the inconsistent rows refer to row data in the baseline data table and the test data table that correspond to the same primary key but have inconsistent data.
[0078] For batch data in non-data table format, the above-mentioned batch data processing unit can first convert the non-data table type batch data into data in data table format through a pre-set adapter, and then perform data consistency verification based on the above-mentioned data consistency verification method for batch data in data table format.
[0079] For streaming data, such as Kafuka data, the streaming data processing unit can first perform streaming integration on the streaming data, convert it into batch data, and then use the data consistency verification method of batch data to perform data consistency verification on the streaming data. For example, for a target node on the offline data link of the recommendation platform, the streaming data processing unit can first write (dump) the streaming data corresponding to the target node into a Hive table, thereby obtaining a test data table and a baseline data table corresponding to the target node in a data table format; then, the test data table corresponding to the read target node and the benchmark data table corresponding to the target node are merged at the row level through a merge operation based on the primary key; the specific fields corresponding to the same primary key are compared respectively, so as to obtain row-level difference data items between the test data table corresponding to the target node and the benchmark data table corresponding to the target node; and finally, the row-level difference data items are combined into the difference data corresponding to the target node.
[0080] For key-value data, it is also divided into two cases: offline key-value data consistency verification and online key-value data consistency verification. For the data consistency verification of offline key-value data, the above-mentioned data consistency verification scheme can also be referred to. That is, the above-mentioned key-value data processing unit can first convert the key-value data and convert it into batch data in the form of a data table. For example, it writes it into a data table, and then uses the data consistency verification method for batch data to perform data consistency verification on the streaming data. For the data consistency verification of online key-value data, in order to ensure online stability, the above-mentioned key-value data processing unit can first sample the key set of baseline data corresponding to a target node on the offline data link of the recommendation platform, and obtain the sampled key set corresponding to the target node. Among them, the above-mentioned key-value data processing unit can determine the sampling rate for sampling the key set of baseline data based on the sampling rate in the configuration information from the user. Then, the key-value data processing unit may perform the following verification operations for each sampling key in the sampling key set corresponding to the target node: perform a point query operation on the test data corresponding to the target node, and obtain the sampling test value corresponding to the sampling key from the test data corresponding to the target node; perform a point query operation on the baseline data corresponding to the target node, and obtain the sampling baseline value corresponding to the sampling key from the baseline data corresponding to the target node; compare the sampling test value with the sampling baseline value; in response to determining that the sampling test value and the sampling baseline value are consistent, the data consistency verification is successful; and in response to determining that the sampling test value and the sampling baseline value are inconsistent, record the sampling key, the sampling test value and the sampling baseline value as a difference data item in the difference data corresponding to the target node.
[0081] From this, it can be seen that, in general, the data verification module 130 in the above-mentioned recommendation platform data verification system can use existing tools to first convert the test data and baseline data corresponding to each node into the format of a data table, and then use data table merging and primary key comparison to find row-level difference data items, thereby obtaining the difference data corresponding to each node that reflects the difference between the test data and the baseline data.
[0082] As mentioned above, the offline data link of the recommendation platform usually involves multiple nodes, and these nodes have the characteristics of front-to-back series connection. The difference in upstream data will be transmitted to the downstream. Therefore, after determining the difference between the test data and the baseline data, attribution analysis is performed on the difference data. Determining whether the difference data is generated by the current node or transmitted from the upstream node is the main goal of the data analysis described in the embodiment of the present disclosure.
[0083] Figure 3 The relationship between the baseline data, test data, and difference data described in the embodiment of the present disclosure and the nodes on the offline data link of the recommendation platform is shown. Figure 3 As shown, baseline data 1 and test data 1 are data generated after node 1 on the offline data link L of the recommendation platform; baseline data 2 and test data 2 are data generated after node 2 on the offline data link L of the recommendation platform; baseline data 3 and test data 3 are data generated after node 3 on the offline data link L of the recommendation platform... In other words, the above baseline data 1 and test data 1 both correspond to node 1 of the offline data link L of the recommendation platform; the above baseline data 2 and test data 2 both correspond to node 2 of the offline data link L of the recommendation platform; the above baseline data 3 and test data 3 both correspond to node 3 of the offline data link L of the recommendation platform. It can be understood that the difference data 1 obtained by performing a data consistency check on the test data 1 based on the baseline data 1 also corresponds to node 1 of the offline data link L of the recommendation platform; the difference data 2 obtained by performing a data consistency check on the test data 2 based on the baseline data 2 also corresponds to node 2 of the offline data link L of the recommendation platform; the difference data 3 obtained by performing a data consistency check on the test data 3 based on the baseline data 3 also corresponds to node 3 of the offline data link L of the recommendation platform. At the same time, it can be understood that the above-mentioned difference data 1, difference data 2 and difference data 3 can each include multiple difference data items, for example, multiple row-level difference data items. In this way, the difference data attribution described in the embodiment of the present disclosure is to determine, for multiple difference data items corresponding to a certain node, whether each difference data item is generated by the current node or its upstream node.
[0084] Based on this, in an embodiment of the present disclosure, the difference data analysis module 140 may include: a difference data attribution unit, mainly used to determine the node generating the difference data. Specifically, in an embodiment of the present disclosure, the difference data attribution unit may also implement the difference data attribution in a variety of ways.
[0085] In some embodiments of the present disclosure, the above-mentioned difference data attribution unit can attribute the difference data based on the primary key set of the difference data. As described above, a piece of difference data can correspond to a certain node on the offline data link of the recommendation platform, and a piece of difference data can include multiple difference data items. Taking the test data and baseline data in the form of a data table as an example, a piece of difference data can include multiple row-level difference data items. In this case, the above-mentioned difference data attribution unit will extract the primary key set of the difference data corresponding to each node respectively. Then, the above-mentioned difference data attribution unit will take each node as the target node in turn according to the order of the nodes on the offline data link of the recommendation platform, and perform the following operations respectively: compare the first primary key set of the difference data corresponding to the target node with the second primary key set of the difference data of the previous level node of the corresponding target node, determine the target primary key that exists in the first primary key set but does not exist in the second primary key set; and attribute the difference data item corresponding to the target primary key in the difference data of the corresponding target node to the target node.
[0086] The difference data attribution process of the difference data attribution unit will be described in detail below with reference to the accompanying drawings and a specific example. Figure 4 An example of difference data corresponding to multiple nodes on the offline data link of the recommendation platform according to the embodiment of the present disclosure is shown. Figure 4As shown, difference data 1 corresponds to node 1 of the recommendation platform offline data link L; difference data 2 corresponds to node 2 of the recommendation platform offline data link L; difference data 3 corresponds to node 3 of the recommendation platform offline data link L. Among them, difference data 1 includes two row-level difference data items:<k1,diff1> ,<k2,diff2> Difference data 2 includes 4 row-level difference data items:<k1,diff1> ,<k2,diff2> ,<k3,diff3> and<k4,diff4> Difference data 3 includes 5 row-level difference data items:<k1,diff1> ,<k2,diff2> ,<k3,diff3> ,<k4,diff4> and<k5,diff5> . Among them, k1, k2, k3, k4 and k5 represent the primary keys of the row-level difference data items, for example, row key. In this case, the above-mentioned difference data attribution unit will first extract the primary key sets of the difference data corresponding to each node respectively. Specifically, the primary key set of the difference data corresponding to node 1 is {k1, k2}; the primary key set of the difference data corresponding to node 2 is {k1, k2, k3, k4}; the primary key set of the difference data corresponding to node 3 is {k1, k2, k3, k4, k5}. Next, for node 1, its corresponding primary key set of difference data {k1, k2} can be compared with the primary key set of difference data corresponding to its previous level node (actually an empty set), and the target primary keys are k1 and k2, then the row-level difference data items with primary keys k1 and k2 in difference data 1 can be attributed to node 1. For node 2, the primary key set of its corresponding difference data {k1, k2, k3, k4} can be compared with the primary key set of the difference data corresponding to its previous level node 1 {k1, k2}, and the target primary keys are k3 and k4. Then, the row-level difference data items with primary keys k3 and k4 in difference data 2 can be attributed to node 2; and the row-level difference data items with primary keys k1 and k2 in difference data 2 have been attributed to the predecessor node 1 of node 2 in the previous round of analysis. Similarly, for node 3, the primary key set of its corresponding difference data {k1, k2, k3, k4, k5} can be compared with the primary key set of the difference data corresponding to its previous level node 2 {k1, k2, k3, k4}, and the target primary key is k5. Then the row-level difference data items with primary key k5 in difference data 3 can be attributed to node 3; and the row-level difference data items with primary keys k3 and k4 in difference data 3 have been attributed to the predecessor node 2 of node 3 in the previous round of analysis; and the row-level difference data items k1 and k2 in difference data 3 have been attributed to the predecessor node 1 of node 3 in an even earlier analysis.
[0087] It can be seen that the above-mentioned differential data attribution based on the primary key set of differential data can quickly realize the attribution of differential data corresponding to each node. However, for the case where a node in the offline data link of the recommendation platform has multiple upstream data sources, that is, the data production logic is not one-way and the primary keys before and after the node change, the attribution of differential data cannot be realized.
[0088] In some other embodiments of the present disclosure, the above-mentioned difference data attribution unit can also perform difference data attribution based on a large language model (LLM). Specifically, in these embodiments, the above-mentioned difference data attribution unit can take each node as a target node in turn according to the order of the nodes on the offline data link of the recommendation platform, and perform the following operations respectively: generate a prompt of the large language model based on the difference data of the corresponding target node, the difference data of the previous level node of the corresponding target node, and the data production logic of the target node; and input the prompt into the large language model, and the large language model outputs the attribution result of whether each difference data item in the difference data of the corresponding target node is attributed to the target node.
[0089] Figure 5 An example of performing differential data attribution based on a large language model by the differential data attribution unit according to an embodiment of the present disclosure is shown. Figure 5 As shown, the above-mentioned difference data attribution unit can first determine the target nodes in sequence according to the order of the nodes on the offline data link of the recommendation platform. For one of the target nodes, the above-mentioned difference data attribution unit can first obtain the difference data corresponding to its previous level node and the difference data corresponding to the target node, and combine the two difference data to obtain the input data part of the large model prompt. The above-mentioned difference data attribution unit can further generate the content part of the large language model prompt based on the data production logic of the target node. Specifically, the above-mentioned data generation logic can be a structured query language (SQL) statement of the production data of the target node. Further, the standardized instruction part and the output indication part of the large language model prompt can also be pre-set. Further, the above-mentioned difference data attribution unit can generate a large language model prompt based on the above-mentioned input data part, content part, instruction part and output indication part, and input the prompt into the large language model, and finally obtain the attribution result of whether each difference data item in the difference data of the target node output by the large language model is attributed to the target node.
[0090] From this, we can see that the use of a large language model can use the data production logic to determine the transmission mode of differential data, thereby effectively attributing differential data. This implementation solution can still effectively attribute differential data even when the primary key of the upstream and downstream differential data changes.
[0091] In addition, in the embodiment of the present disclosure, the above-mentioned difference data analysis module 140 may further include: a difference data statistical analysis unit, which is mainly used to complete the statistical analysis of the difference data, for example, to determine the difference rate between the test data and the baseline data or the difference rate of each field, etc. Specifically, the above-mentioned difference data statistical analysis unit can be used to analyze whether the structure of the test data and the baseline data is consistent based on the difference data. Specifically, for the test data and the baseline data of the data table type, whether the above-mentioned structure is consistent may include: whether the structure definition (Schema) is consistent and whether the number of rows and columns is consistent, etc. Further, the above-mentioned difference data statistical analysis unit can also analyze the difference rate between the test data and the baseline data based on the difference data. Specifically, for the test data and the baseline data of the data table type, the above-mentioned difference rate can be specifically expressed as the table difference rate and the difference rate of each field. In the embodiment of the present disclosure, the above-mentioned table difference rate can be defined as the ratio of the number of rows in the test table that are consistent with the baseline table to the total number of rows in the baseline table. The difference rate of each field can be obtained by statistically analyzing the row-level difference data items in the difference data.
[0092] It can be seen from this that the recommendation platform data verification system provided by some embodiments of the present disclosure can quickly verify the data of the recommendation platform offline data link through automated data difference detection, discover and locate data differences, significantly improve the efficiency of data verification, and reduce the workload and error probability of manual intervention. Furthermore, the above-mentioned recommendation platform data verification system can effectively eliminate the impact of data differences transmitted from the upstream links through the upstream and downstream data difference attribution mechanism, avoid the limitation of traditional tools that cannot accurately locate the root cause of the difference data, and provide higher quality difference data analysis results. In addition, the above-mentioned recommendation platform data verification system can also support multiple data types and multiple data sources, including batch data, streaming data, key-value data, etc., and can flexibly respond to the complex data flow environment of the recommendation platform and meet the needs of the recommendation platform.
[0093] Corresponding to the above-mentioned recommendation platform data verification system, the embodiment of the present disclosure also discloses a recommendation platform data verification method. Figure 6 The implementation process of the recommendation platform data verification method described in the embodiment of the present disclosure is shown. Figure 6 As shown, the above-mentioned recommendation platform data verification method may include the following steps:
[0094] In step 610, the baseline data corresponding to each node on the recommendation platform offline data link is read based on the storage path of the baseline data.
[0095] In step 620, the test data corresponding to each node on the offline data link of the recommendation platform is read based on the storage path of the test data.
[0096] In the embodiment of the present disclosure, the storage path of the baseline data and the storage path of the test data can be obtained from the configuration information of the data verification task from the user. Moreover, the execution order of the above steps 610 and 620 is not limited by the order of the step numbers.
[0097] In step 630, data consistency check is performed on the test data corresponding to each node based on the baseline data corresponding to each node, and difference data reflecting the difference between the test data and the baseline data corresponding to each node is determined respectively.
[0098] In an embodiment of the present disclosure, the difference data may include at least one difference data item. In addition, for different types of data, multiple methods may be used to determine the difference data.
[0099] Specifically, as mentioned above, for batch data in the form of data tables, at this time, the above test data is referred to as a test data table, and the above baseline data is referred to as a baseline data table. The above step 620 may include: for a certain target node on the offline data link of the recommendation platform, a row-level merge is performed on the test data table corresponding to the target node and the benchmark data table corresponding to the target node through a merge operation based on the primary key; data is compared on the fields corresponding to the same primary key to obtain row-level difference data items between the test data table corresponding to the target node and the benchmark data table corresponding to the target node; and the row-level difference data items are combined into difference data corresponding to the target node. As also mentioned above, the above row-level difference data items may include: missing rows, redundant rows, and inconsistent rows.
[0100] For batch data, streaming data or offline key-value data in non-table format, the above step 620 may include: for a target node on the offline data link of the recommendation platform, converting the baseline data corresponding to the target node into a baseline data table corresponding to the target node, and converting the test data corresponding to the target node into a test data table corresponding to the target node; performing a row-level merge on the test data table corresponding to the target node and the benchmark data table corresponding to the target node through a primary key-based merge operation; performing data comparison on the fields corresponding to the same primary key to obtain row-level difference data items between the test data table corresponding to the target node and the benchmark data table corresponding to the target node; and combining the row-level difference data items into difference data corresponding to the target node. The specific conversion method is completed based on the original data format, for example, it can be implemented using a suitable adapter.
[0101] For the data consistency check of online key-value data, the above step 620 may include: for a certain target node on the offline data link of the recommendation platform, sampling the key set of the baseline data corresponding to the target node to obtain the sampling key set corresponding to the target node; and performing the following verification operations on each sampling key in the sampling key set corresponding to the target node: performing a point query operation on the test data corresponding to the target node, and obtaining the sampling test value corresponding to the sampling key from the test data corresponding to the target node (the value obtained by point querying the test data); performing a point query operation on the baseline data corresponding to the target node, and obtaining the sampling baseline value corresponding to the sampling key from the baseline data corresponding to the target node (the value obtained by point querying the baseline data); comparing the sampling test value with the sampling baseline value; in response to determining that the sampling test value and the sampling baseline value are consistent, the data consistency check is successful; and in response to determining that the sampling test value and the sampling baseline value are inconsistent, recording the sampling key, the sampling test value and the sampling baseline value as a difference data item in the difference data corresponding to the target node.
[0102] In step 640, attribution analysis is performed on the difference data corresponding to each node to determine the analysis result of the difference data corresponding to each node.
[0103] The above analysis results include: nodes on the offline data link of the recommendation platform that generate each difference data item in the difference data.
[0104] As mentioned above, in some embodiments of the present disclosure, the above-mentioned attribution analysis may specifically include: respectively extracting the primary key sets of the difference data corresponding to each node; taking each node as the target node in turn according to the order of the nodes on the offline data link of the recommendation platform, and respectively performing the following operations: comparing the first primary key set of the difference data corresponding to the target node with the second primary key set of the difference data of the previous level node corresponding to the target node, determining the target primary key that exists in the first primary key set but does not exist in the second primary key set; and attributing the difference data item corresponding to the target primary key in the difference data corresponding to the target node to the target node.
[0105] In some other embodiments of the present disclosure, the above-mentioned attribution analysis may specifically include: taking each node as a target node in turn according to the order of the nodes on the offline data link of the recommendation platform, and performing the following operations respectively: generating a prompt of a large language model based on the difference data of the corresponding target node, the difference data of the previous level node of the corresponding target node, and the data production logic of the target node; and inputting the prompt into the large language model, and the large language model outputs the attribution result of whether each difference data item in the difference data of the corresponding target node is attributed to the target node.
[0106] In step 650, the analysis results of the difference data corresponding to each of the above nodes are displayed.
[0107] In an embodiment of the present disclosure, the above-mentioned recommendation platform data verification method may further include: performing statistical analysis on the difference data corresponding to each node respectively, and obtaining the difference rate between the test data and the baseline data corresponding to each node; and adding the difference rate to the analysis result of the difference data corresponding to its corresponding node. Specifically, for the test data and baseline data of the data table type, the above-mentioned difference rate can be specifically expressed as the table difference rate and the difference rate of each field. As mentioned above, the above-mentioned table difference rate can be defined as the ratio of the number of rows in the test table that are consistent with the baseline table to the total number of rows in the baseline table. The difference rate of each field mentioned above can also be obtained by statistically analyzing the row-level difference data items in the difference data.
[0108] It can be seen that the recommendation platform data verification method provided by some embodiments of the present disclosure can quickly verify the data of the offline data link of the recommendation platform through automated data difference troubleshooting, discover and locate data differences, significantly improve the efficiency of data verification, and reduce the workload and error probability of manual intervention. Furthermore, through the upstream and downstream data difference attribution mechanism, the influence of data differences transmitted from the upstream link can be effectively eliminated, avoiding the limitation of traditional tools that cannot accurately locate the root cause of the difference data, and providing higher quality difference data analysis results. In addition, the above-mentioned recommendation platform data verification method can also support multiple data types and multiple data sources, including batch data, streaming data, key-value data, etc., and can flexibly cope with the complex data flow environment of the recommendation platform and meet the needs of the recommendation platform.
[0109] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the recommended platform data verification method described in any of the above embodiments is implemented.
[0110] Figure 7 A more specific hardware structure diagram of an electronic device provided in this embodiment is shown, and the device may include: a processor 2010, a memory 2020, an input / output interface 2030, a communication interface 2040, and a bus 2050. The processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040 are connected to each other in communication within the device through the bus 2050.
[0111] The processor 2010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0112] The memory 2020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 2020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 2020 and are called and executed by the processor 2010.
[0113] The input / output interface 2030 is used to connect input / output devices to realize information input and output. The input / output devices can be configured in the device as components, or can be externally connected to the device to provide corresponding functions. The input devices can include microphones, various sensors, etc., and the output devices can include displays, speakers, vibrators, indicator lights, etc.
[0114] The communication interface 2040 is used to connect a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0115] The bus 2050 includes a path that transmits information between the various components of the device (eg, the processor 2010, the memory 2020, the input / output interface 2030, and the communication interface 2040).
[0116] It should be noted that, although the above device only shows the processor 2010, the memory 2020, the input / output interface 2030, the communication interface 2040, and the bus 2050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include the components necessary for implementing the embodiments of the present specification, and does not necessarily include all the components shown in the figure.
[0117] The electronic device of the above embodiment is used to implement the corresponding recommendation platform data verification method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0118] Based on the same inventive concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the recommendation platform data verification method described in any of the above embodiments.
[0119] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0120] The computer instructions stored in the storage medium of the above embodiments are used to enable the computer to execute the task processing method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0121] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.
[0122] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0123] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0124] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A method for verifying recommended platform data, comprising: Read the baseline data corresponding to each node on the offline data link of the recommendation platform based on the storage path of the baseline data; Read the test data corresponding to each node based on the storage path of the test data; Performing data consistency check on the test data corresponding to each node based on the baseline data corresponding to each node, and determining the difference data reflecting the difference between the test data and the baseline data corresponding to each node; wherein the difference data includes at least one difference data item; Performing attribution analysis on the difference data corresponding to each node to determine the analysis results of the difference data corresponding to each node; wherein the analysis results include: the node on the offline data link of the recommendation platform that generates each difference data item in the difference data; and The analysis results of the difference data corresponding to each node are displayed.
2. The method according to claim 1, wherein: The baseline data corresponding to each node is a baseline data table; the test data corresponding to each node is a test data table; and the data consistency check of the test data corresponding to each node based on the baseline data corresponding to each node includes: For a target node on the offline data link of the recommendation platform, a test data table corresponding to the target node and a benchmark data table corresponding to the target node are merged at the row level through a primary key-based merge operation; Comparing data on fields corresponding to the same primary key to obtain row-level difference data items between the test data table corresponding to the target node and the benchmark data table corresponding to the target node; and The row-level difference data items are combined into difference data corresponding to the target node.
3. The method according to claim 1, wherein: The baseline data corresponding to each node is batch data, streaming data or offline key-value data in a non-data table format; the test data corresponding to each node is batch data, streaming data or offline key-value data in a non-data table format; and the data consistency check of the test data corresponding to each node based on the baseline data corresponding to each node includes: For a target node on the offline data link of the recommendation platform, converting the baseline data corresponding to the target node into a baseline data table corresponding to the target node, and converting the test data corresponding to the target node into a test data table corresponding to the target node; Performing a row-level merge on the test data table corresponding to the target node and the benchmark data table corresponding to the target node through a primary key-based merge operation; Comparing data on fields corresponding to the same primary key to obtain row-level difference data items between the test data table corresponding to the target node and the benchmark data table corresponding to the target node; and The row-level difference data items are combined into difference data corresponding to the target node.
4. The method according to claim 1, wherein: The baseline data corresponding to each node is online key-value data; the test data corresponding to each node is online key-value data; and the data consistency check of the test data corresponding to each node based on the baseline data corresponding to each node includes: For a target node on the offline data link of the recommendation platform, sampling a key set of baseline data corresponding to the target node to obtain a sampled key set corresponding to the target node; and The following verification operations are performed for each sampling key in the sampling key set corresponding to the target node: Performing a point query operation on the test data corresponding to the target node, and obtaining a sampled test value corresponding to the sampling key from the test data corresponding to the target node; Performing a point query operation on the baseline data corresponding to the target node, and obtaining a sampling baseline value corresponding to the sampling key from the baseline data corresponding to the target node; comparing the sampled test value with the sampled baseline value; In response to determining that the sampled test value and the sampled baseline value are consistent, the data consistency check succeeds; and In response to determining that the sampled test value and the sampled baseline value are inconsistent, the sampling key, the sampled test value and the sampled baseline value are recorded as a difference data item in the difference data corresponding to the target node.
5. The method according to claim 1, wherein: The performing attribution analysis on the difference data corresponding to each node includes: Extracting primary key sets of difference data corresponding to each node respectively; According to the order of nodes on the offline data link of the recommendation platform, each node is taken as the target node in turn, and the following operations are performed respectively: Compare a first primary key set corresponding to the difference data of the target node with a second primary key set corresponding to the difference data of the node one level before the target node, and determine a target primary key that exists in the first primary key set but does not exist in the second primary key set; and The difference data item corresponding to the target primary key in the difference data corresponding to the target node is attributed to the target node.
6. The method according to claim 1, wherein: The performing attribution analysis on the difference data corresponding to each node includes: According to the order of nodes on the offline data link of the recommendation platform, each node is taken as the target node in turn, and the following operations are performed respectively: Generate a prompt of a large language model based on the difference data corresponding to the target node, the difference data corresponding to the node one level before the target node, and the data production logic of the target node; and The prompt is input into a large language model, and the large language model outputs an attribution result of whether each difference data item in the difference data corresponding to the target node is attributed to the target node.
7. The method according to claim 1, further comprising: Performing statistical analysis on the difference data corresponding to each node to obtain the difference rate between the test data and the baseline data corresponding to each node; as well as The difference rate is added to the analysis result of the difference data corresponding to its corresponding node.
8. A recommendation platform data verification system, comprising: A data storage module, used to store baseline data corresponding to each node on the offline data link of the recommendation platform, test data corresponding to each node, and difference data corresponding to each node reflecting the difference between the test data and the baseline data; A configuration module, used to receive configuration information of a data verification task; wherein the configuration information includes: a storage path of baseline data, a storage path of test data, and a storage path of difference data; a data verification module, configured to read the baseline data corresponding to each node from the data storage module based on the storage path of the baseline data, read the test data corresponding to each node from the data storage module based on the storage path of the test data, perform data consistency verification on the test data corresponding to each node based on the baseline data corresponding to each node, respectively determine the difference data corresponding to each node, and store the difference data corresponding to each node in the data storage module; wherein the difference data includes at least one difference data item; a difference data analysis module, configured to read the difference data corresponding to each node from the data storage module, perform attribution analysis on the difference data corresponding to each node, and determine the analysis results of the difference data corresponding to each node; wherein the analysis results include: a node on the offline data link of the recommendation platform that generates each difference data item in the difference data; and The analysis result display module is used to display the analysis results of the difference data corresponding to each node.
9. The system according to claim 8, further comprising: The scheduling module is used to pass the configuration information to the data verification module, schedule the data verification module to perform the data consistency verification task, and after the data verification module is executed, pass the configuration information to the difference data analysis module, and schedule the difference data analysis module to perform the difference data analysis task.
10. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for verifying recommended platform data as described in any one of claims 1 to 7 is implemented.
11. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the recommendation platform data verification method according to any one of claims 1 to 7.
12. A computer program product, comprising computer program instructions, which, when executed on a computer, enable the computer to execute the recommendation platform data verification method according to any one of claims 1 to 7.