Data quality identification method, device, computer equipment and storage medium
By linear regression processing of survey data from multiple data platforms and obtaining quality evaluation parameters, the problem of low accuracy of data quality recognition caused by manual subjective evaluation is solved, and more accurate data quality recognition is achieved.
Patent Information
- Application Number
- CN202110250433.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-08
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-03-08
AI Technical Summary
In the questionnaire scenario, in the prior art, manual subjective evaluation of data quality is prone to errors, resulting in low accuracy of data quality identification.
By obtaining survey data collected by multiple data platforms, performing linear regression processing, obtaining quality evaluation parameters, determining the quality identification results of the data platform, and finally determining the quality identification results of the target survey data.
It improves the accuracy of data quality identification, avoids errors caused by manual subjective analysis, and ensures the accuracy of data quality identification.
Smart Images

Figure CN112862355B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data quality identification method, apparatus, computer equipment, and storage medium. Background Art
[0002] With the development of big data technology, it is usually necessary to perform data analysis on the collected data; however, before performing data analysis, it is necessary to evaluate the data quality of the collected data to ensure the reliability of the data analysis results.
[0003] In questionnaire survey scenarios, the data quality of survey data is generally evaluated manually and subjectively. However, errors are prone to occur during the process of manual subjective evaluation of data quality, resulting in low accuracy in identifying data quality. Summary of the Invention
[0004] Based on this, it is necessary to provide a data quality identification method, device, computer equipment and storage medium that can improve the recognition accuracy of data quality in response to the above technical problems.
[0005] A data quality identification method, comprising:
[0006] Obtaining target survey data; the target survey data includes survey data collected from at least two data platforms; the survey data of each data platform is collected from the same user group using the same questionnaire by the corresponding data platform, including survey indicators and statistical data corresponding to the survey indicators;
[0007] According to the target survey data, a statistical data set of each survey indicator of the two-two data platform is obtained;
[0008] Performing linear regression processing on the statistical data set of each survey indicator of the pairwise data platform to obtain multiple quality assessment parameters of the survey data of the pairwise data platform;
[0009] Determining a quality identification result of the survey data of the pairwise data platform according to a plurality of quality assessment parameters of the survey data of the pairwise data platform;
[0010] According to the quality identification results of the survey data of the pairwise data platforms, a target quality identification result of the target survey data is determined.
[0011] A data quality identification device, comprising:
[0012] A data acquisition module is configured to acquire target survey data; the target survey data includes survey data collected via at least two data platforms; the survey data on each data platform is collected from the same user group using the same questionnaire, including survey indicators and statistical data corresponding to the survey indicators;
[0013] A set acquisition module is used to obtain a statistical data set of each survey indicator of the two-by-two data platform according to the target survey data;
[0014] a parameter acquisition module, configured to perform linear regression processing on a statistical data set of each survey indicator of the pairwise data platform to obtain a plurality of quality assessment parameters of the survey data of the pairwise data platform;
[0015] a result determination module, configured to determine a quality identification result of the survey data of the pairwise data platform according to a plurality of quality assessment parameters of the survey data of the pairwise data platform;
[0016] The quality identification module is used to determine the target quality identification result of the target survey data based on the quality identification result of the survey data of the pairwise data platform.
[0017] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0018] Obtaining target survey data; the target survey data includes survey data collected from at least two data platforms; the survey data of each data platform is collected from the same user group using the same questionnaire by the corresponding data platform, including survey indicators and statistical data corresponding to the survey indicators;
[0019] According to the target survey data, a statistical data set of each survey indicator of the two-two data platform is obtained;
[0020] Performing linear regression processing on the statistical data set of each survey indicator of the pairwise data platform to obtain multiple quality assessment parameters of the survey data of the pairwise data platform;
[0021] Determining a quality identification result of the survey data of the pairwise data platform according to a plurality of quality assessment parameters of the survey data of the pairwise data platform;
[0022] According to the quality identification results of the survey data of the pairwise data platforms, a target quality identification result of the target survey data is determined.
[0023] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0024] Obtaining target survey data; the target survey data includes survey data collected from at least two data platforms; the survey data of each data platform is collected from the same user group using the same questionnaire by the corresponding data platform, including survey indicators and statistical data corresponding to the survey indicators;
[0025] According to the target survey data, a statistical data set of each survey indicator of the two-two data platform is obtained;
[0026] Performing linear regression processing on the statistical data set of each survey indicator of the pairwise data platform to obtain multiple quality assessment parameters of the survey data of the pairwise data platform;
[0027] Determining a quality identification result of the survey data of the pairwise data platform according to a plurality of quality assessment parameters of the survey data of the pairwise data platform;
[0028] According to the quality identification results of the survey data of the pairwise data platforms, a target quality identification result of the target survey data is determined.
[0029] The above-mentioned data quality identification method, device, computer equipment and storage medium obtain target survey data; the target survey data includes survey data collected by at least two data platforms; the survey data of each data platform is collected by the corresponding data platform for the same user group based on the same questionnaire, including survey indicators and statistical data corresponding to the survey indicators; then, based on the target survey data, a statistical data set of each survey indicator of each data platform is obtained; linear regression processing is performed on the statistical data set of each survey indicator of each data platform to obtain multiple quality assessment parameters of the survey data of each data platform; then, based on the multiple quality assessment parameters of the survey data of the each data platform, the quality identification results of the survey data of the each data platform are determined; finally, based on the quality identification results of the survey data of the each data platform, the target quality identification results of the target survey data are determined; in this way, the purpose of determining the target quality identification results of the target survey data based on the quality identification results of the survey data of the each data platform is achieved, and the survey data collected by multiple data platforms are comprehensively considered, so that the quality identification of the survey data is more accurate, thereby improving the recognition accuracy of data quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A schematic diagram of the structure of a distributed system provided in one embodiment applied to a blockchain system;
[0031] Figure 2 A schematic diagram of a block structure provided in one embodiment;
[0032] Figure 3 This is a diagram of an application environment of a data quality identification method in one embodiment;
[0033] Figure 4 1 is a flow chart of a data quality identification method according to an embodiment;
[0034] Figure 5 A flowchart illustrating steps for obtaining multiple quality assessment parameters of survey data on a pairwise data platform in one embodiment;
[0035] Figure 6 A flowchart illustrating steps for determining target quality identification results for target survey data in one embodiment;
[0036] Figure 7 A flowchart of the steps of obtaining new target survey data in one embodiment;
[0037] Figure 8 is a flow chart of a data quality identification method according to another embodiment;
[0038] Figure 9 A distribution comparison diagram of dual-channel data in one embodiment;
[0039] Figure 10 2. A distribution diagram of the difference between channel B and channel A in one embodiment;
[0040] Figure 11 is a structural block diagram of a data quality identification device in one embodiment;
[0041] Figure 12 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0043] The data quality identification system involved in the embodiment of the present application can be a distributed system formed by connecting multiple nodes (any form of computing devices in the access network, such as servers, data platforms) through network communication.
[0044] Taking the distributed system as the blockchain system as an example, see Figure 1 , Figure 1This is an optional structural diagram of the distributed system 100 provided in an embodiment of the present application applied to a blockchain system. It is composed of multiple nodes 200 (any form of computing device connected to the network, such as a server or data platform). Nodes 200 form a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol that runs on top of the Transmission Control Protocol (TCP). In a distributed system, any machine, such as a server or data platform, can join and become a node 200. Node 200 includes a hardware layer, an intermediate layer, an operating system layer, and an application layer.
[0045] See also Figure 1 The functions of each node 200 in the blockchain system shown include:
[0046] (1) Routing: This is a basic function of the node 200, used to support communication between nodes.
[0047] In addition to the routing function, node 200 may also have the following functions:
[0048] (2) Application, which is used to be deployed in the blockchain to implement specific business according to actual business needs, record data related to the implementation of the function to form record data, carry a digital signature in the record data to indicate the source of the task data, and send the record data to other nodes 200 in the blockchain system, so that other nodes 200 can add the record data to the temporary block when they successfully verify the source and integrity of the record data.
[0049] (3) Blockchain, including a series of blocks that are connected to each other in the order of their generation. Once a new block is added to the blockchain, it will not be removed. The block records the record data submitted by node 200 in the blockchain system.
[0050] See also Figure 2 , Figure 2 This is an optional schematic diagram of the block structure provided by an embodiment of the present application. Each block includes the hash value of the transaction record stored in the block (the hash value of the current block) and the hash value of the previous block. The blocks are connected by hash values to form a blockchain. In addition, the block may also include information such as the timestamp when the block was generated. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains relevant information used to verify the validity of its information (anti-counterfeiting) and generate the next block.
[0051] Furthermore, the data quality identification method provided by this application can also be applied to Figure 3 In the application environment shown, the application environment includes a data quality identification system, which can be a blockchain system. Specifically, refer to Figure 3 The data quality identification system includes a server 302 and multiple data platforms 304 (e.g., data platform 304a, data platform 304b, etc.). The server 302 communicates with the multiple data platforms 304 via a network. In a questionnaire survey scenario, the server 302 obtains target survey data; the target survey data includes survey data collected by at least two data platforms 304; the survey data of each data platform 304 is collected by the corresponding data platform for the same user group based on the same questionnaire, including survey indicators and statistical data corresponding to the survey indicators; based on the target survey data, a statistical data set of each survey indicator of each data platform is obtained; linear regression processing is performed on the statistical data set of each survey indicator of each data platform to obtain multiple quality assessment parameters of the survey data of each data platform; based on the multiple quality assessment parameters of the survey data of each data platform, a quality identification result for the survey data of each data platform is determined; based on the quality identification result of the survey data of each data platform, a target quality identification result for the target survey data is determined.
[0052] Among them, server 302 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Data platform 304 refers to a platform for collecting user data, which can be a smart phone, tablet computer, laptop computer, desktop computer, etc., but is not limited to this. The data platform and the server can be directly or indirectly connected through wired or wireless communication, and this application does not limit this.
[0053] In one embodiment, Figure 4 As shown, a data quality identification method is provided, which is applied to Figure 3 The following steps are used as an example to illustrate the server in the example:
[0054] Step S402, obtaining target survey data; the target survey data includes survey data collected by at least two data platforms; the survey data of each data platform is collected by the corresponding data platform for the same user group based on the same questionnaire, including survey indicators and statistical data corresponding to the survey indicators.
[0055] Among them, target survey data refers to survey data collected from two or more data platforms, such as survey data collected through multiple channels.
[0056] Data platforms, such as SMS and messaging platforms, distribute surveys to users and collect survey data. Each data platform distributes the same survey to the same user group. The survey includes multiple questions, each of which includes multiple indicators. For example, the question "Which W Pay features have you used in the past year?" corresponds to the following indicators: transferring money to friends, transferring money to bank cards, transferring money to mobile numbers, and paying with QR codes. Statistics refer to the percentage of users corresponding to each indicator. For example, the percentage of users transferring money to friends is 50%, the percentage of users transferring money to bank cards is 30%, and the percentage of users transferring money to mobile numbers is 10%.
[0057] Among them, user groups refer to groups that need to be surveyed, such as netizens, students, mothers and babies, female user groups aged 18-35, etc.
[0058] Specifically, the server obtains survey data collected from at least two data platforms as target survey data. For example, the server obtains the URL of a survey form and then sends the URL to different data platforms. Each data platform then sends the URL to a designated user through a targeted survey, prompting the user to complete the survey form and then collects and saves the corresponding survey data. Finally, the server obtains the survey data collected by each data platform as target survey data.
[0059] It should be noted that each data platform is an independent data platform, and data is collected from the same user group based on the same questionnaire. The ideal situation is that the survey data collected by each data platform in the target survey data are completely consistent, which indirectly indicates that the data quality of the target survey data is relatively high; but in reality, the survey data collected by each data platform in the target survey data cannot be completely consistent. Therefore, through the data quality identification method of this application, it is possible to quantitatively judge whether the data quality of the survey data collected by each data platform in the target survey data is consistent. If they are different from the same source, it is determined that the data quality of the target survey data meets the standards; in this way, the survey data collected by multiple data platforms are comprehensively considered, which is conducive to improving the recognition accuracy of data quality and avoiding the defect that errors are prone to occur in the data quality of survey data through manual subjective analysis, resulting in low data quality recognition accuracy.
[0060] Step S404: Obtain a statistical data set of each survey indicator of the pairwise data platform based on the target survey data.
[0061] Among them, the statistical data set refers to the set of statistical data of the same survey indicator in two data platforms; for example, in data platform A, in the question "Which functions of W payment have you used in the past year?", the proportion of users who transferred money to friends was 50%, and the proportion of users who transferred money to mobile phone numbers was 60%; in data platform B, in the question "Which functions of W payment have you used in the past year?", the proportion of users who transferred money to friends was 48%, and the proportion of users who transferred money to mobile phone numbers was 59%. Therefore, in data platforms A and B, the statistical data set of transfers to friends is (50%, 48%), and the statistical data set of transfers to mobile phone numbers is (60%, 59%).
[0062] Specifically, the server obtains statistical data of each survey indicator of each data platform based on the target survey data; and obtains a set of statistical data of each survey indicator of each data platform based on the statistical data of each survey indicator.
[0063] Step S406 , performing linear regression processing on the statistical data set of each survey indicator of the pairwise data platform to obtain multiple quality assessment parameters of the survey data of the pairwise data platform.
[0064] Among them, quality assessment parameters refer to parameters used to evaluate the data quality of survey data, such as the coefficient and intercept of a linear equation.
[0065] Specifically, the server establishes a mapping relationship between the statistical data of the two data platforms through linear regression; based on the mapping relationship between the statistical data of the two data platforms, multiple quality assessment parameters of the survey data of the two data platforms are obtained.
[0066] Step S408 : determining a quality identification result of the survey data of the two-by-two data platform according to a plurality of quality evaluation parameters of the survey data of the two-by-two data platform.
[0067] Among them, under normal circumstances, since each data platform is independent of each other and collects data from the same user group based on the same questionnaire, the data quality of the survey data from different data platforms is homogeneous, that is, the same quality from different sources, which means that the data quality of the survey data from different data platforms meets the standards; under abnormal circumstances, the data quality of the survey data from different data platforms is different, that is, different quality from different sources, which means that the data quality of the survey data from different data platforms does not meet the standards.
[0068] Specifically, the server compares multiple quality assessment parameters of the survey data on the two-by-two data platform with corresponding thresholds to obtain a comparison result. Based on the comparison result, the server determines the quality identification result for the survey data on the two-by-two data platform. For example, if multiple quality assessment parameters of the survey data on the two-by-two data platform all meet the corresponding thresholds, then the data quality of the survey data on the two-by-two data platform is determined to meet the data quality standards.
[0069] Step S410: determining a target quality identification result for target survey data based on the quality identification results for the survey data of the pairwise data platforms.
[0070] Among them, under normal circumstances, the data quality of the survey data of all data platforms is homogeneous, that is, the data quality of different sources is homogeneous, which means that the data quality of the survey data of different data platforms meets the standards, and further indicates that the data quality of the target survey data meets the standards; under abnormal circumstances, the data quality of the survey data of all data platforms is different, that is, different sources are different, which means that the data quality of the target survey data does not meet the standards.
[0071] For example, if the data quality of the survey data of all pairwise data platforms meets the standards, it is determined that the data quality of the target survey data meets the standards.
[0072] In the above data quality identification method, target survey data is obtained; the target survey data includes survey data collected by at least two data platforms; the survey data of each data platform is collected by the corresponding data platform for the same user group based on the same questionnaire, including survey indicators and statistical data corresponding to the survey indicators; then, based on the target survey data, a set of statistical data of each survey indicator of each data platform is obtained; linear regression processing is performed on the statistical data set of each survey indicator of each data platform to obtain multiple quality assessment parameters of the survey data of each data platform; then, based on the multiple quality assessment parameters of the survey data of the each data platform, the quality identification results of the survey data of the each data platform are determined; finally, based on the quality identification results of the survey data of the each data platform, the target quality identification results of the target survey data are determined; in this way, the purpose of determining the target quality identification results of the target survey data based on the quality identification results of the survey data of the each data platform is achieved, and the survey data collected by multiple data platforms are comprehensively considered, so that the quality identification of the survey data is more accurate, thereby improving the recognition accuracy of data quality.
[0073] In one embodiment, the above step S404, before obtaining the statistical data set of each survey indicator of the pairwise data platform based on the target survey data, further includes: filtering invalid data in the target survey data to obtain filtered target survey data.
[0074] Invalid data may refer to data that has exceeded the time limit for filling in the answer, or data that is inconsistent with the data filled in by the user.
[0075] Specifically, the server obtains a preset invalid data filtering instruction, and filters the invalid data in the target survey data according to the preset invalid data filtering instruction to obtain filtered target survey data; wherein the preset invalid data filtering instruction is an instruction for filtering out invalid data in the survey data.
[0076] For example, assuming that the prescribed filling time is one minute and the user fills in two minutes, the survey data corresponding to the user is invalid data.
[0077] Furthermore, the server can also use weight control to adjust the unreasonable distribution of some key attributes of the recovered survey data; for example, the expected male-female ratio of the recovered data is 60%:40%, but the actual recovered ratio is 50%:50%. At this time, a weighted questionnaire processing operation is required.
[0078] Then, the above step S404, obtaining a statistical data set of each survey indicator of the two-by-two data platform according to the target survey data, includes: obtaining a statistical data set of each survey indicator of the two-by-two data platform according to the filtered target survey data.
[0079] Specifically, the server obtains statistical data of each survey indicator of each data platform based on the filtered target survey data, and further obtains a set of statistical data of each survey indicator of each data platform.
[0080] In this embodiment, invalid data in the target survey data is first filtered, and then a statistical data set of each survey indicator of the pairwise data platform is obtained based on the filtered target survey data, so that the data quality identification of subsequent target survey data is more accurate, and the recognition accuracy of data quality is further improved.
[0081] In one embodiment, Figure 5 As shown, the above step S406, performing linear regression processing on the statistical data set of each survey indicator of the two-by-two data platform to obtain multiple quality assessment parameters of the survey data of the two-by-two data platform, specifically includes the following steps:
[0082] Step S502 : performing linear regression processing on the statistical data sets of the various survey indicators of the two data platforms to obtain a mapping relationship between the statistical data of the two data platforms.
[0083] Step S504: Based on the statistical data set and mapping relationship of each survey indicator of the two-by-two data platform, the regression fitting parameters of multiple preset dimensions of the survey data of the two-by-two data platform are obtained, and the regression fitting parameters of the multiple preset dimensions are used as multiple quality evaluation parameters of the survey data of the two-by-two data platform.
[0084] The regression fitting parameters of the preset dimensions refer to the coefficient, intercept, determination coefficient in the linear equation and the proportion of the absolute value of the difference between the statistical data of different data platforms less than 10%; assuming that the linear equation is y=ax+b, the coefficient is a, the intercept is b, and the determination coefficient is R 2 , which is used to represent the effect of linear regression, specifically the ratio of the estimated value of y to the actual value, ranging from 0 to 1; if the determination coefficient is 1, the statistical data of the two data platforms have a good correlation, and there is no difference between the estimated value of y and the actual value; on the contrary, if the determination coefficient is 0, the regression formula y = ax + b cannot be used to predict the value of y.
[0085] Among them, if the coefficient is less than 0.9 or greater than 1.1, the quality assessment does not meet the standard; if the intercept is less than -0.1 or greater than 0.1, the quality assessment does not meet the standard; if the determination coefficient is less than 0.6, it means that the regression effect is not good and the quality assessment does not meet the standard; if the absolute value of the difference in statistical data from different data platforms is less than 10% and the proportion is less than 70%, the quality assessment does not meet the standard.
[0086] It should be noted that, assuming that in the question "Which functions of W Pay have you used in the past year", the statistical data set of transfers to friends in data platforms A and B is (50%, 48%), and the statistical data set of transfers to mobile phone numbers is (60%, 59%), then the absolute value of the difference in the corresponding statistical data is 2% and 1%.
[0087] Specifically, the server constructs a mapping relationship between the statistical data of each data platform through linear regression to obtain a linear equation. For example, assuming that the statistical data sets for data platforms A and B are (59%, 60%), (55%, 56%), (34%, 35%), and (62%, 63%), and 60% = 1 × 59% + 1%, 56% = 1 × 55% + 1%, 35% = 1 × 34% + 1%, and 63% = 1 × 62% + 1%, it means that the mapping relationship between the statistical data of data platforms A and B is y = 1 × x + 1%, which satisfies the linear equation expression, indicating that the final mapping relationship is a linear equation y = 1 × x + 1%. Then, according to the linear equation, the coefficient, intercept and determination coefficient are obtained. For example, for the above linear equation y = 1 × x + 1%, the coefficient is 1 and the intercept is 1%; according to the statistical data set of each survey indicator of the two-two data platform, the absolute value of the difference of the statistical data of each survey indicator of the two-two data platform is obtained; for example, for the statistical data sets (59%, 60%), (55%, 56%), (34%, 35%), and (62%, 63%), the absolute value of the difference of the statistical data of each survey indicator of the A and B data platforms is 1%; then the proportion of the absolute value of the difference of the statistical data of each survey indicator of the two-two data platform that is less than the preset threshold is obtained, and finally the proportion of the absolute value of the difference of the statistical data of each survey indicator of the two-two data platform that is less than the preset threshold, the coefficient, the intercept and the determination coefficient in the linear equation are used as the regression fitting parameters of multiple preset dimensions of the survey data of the two-two data platform, and the regression fitting parameters of the multiple preset dimensions are used as multiple quality assessment parameters of the survey data of the two-two data platform.
[0088] In this embodiment, obtaining multiple quality assessment parameters of the survey data of the two-by-two data platform facilitates determining the quality identification result of the survey data of the two-by-two data platform based on the multiple quality assessment parameters of the survey data of the two-by-two data platform.
[0089] In one embodiment, the above-mentioned step S408 determines the quality identification results of the survey data of the two data platforms based on multiple quality assessment parameters of the survey data of the two data platforms, including: if the multiple quality assessment parameters of the survey data of the two data platforms all meet the corresponding thresholds, then it is determined that the quality identification results of the survey data of the two data platforms are the same quality level.
[0090] Among them, the quality identification result of the survey data of the two data platforms is the same quality level, which means that the quality level of the survey data of the two data platforms is the same, specifically means that the data quality of the survey data of the two data platforms is homogeneous, that is, the data quality of the survey data of the two data platforms meets the standards.
[0091] Among them, each quality assessment parameter is set with a corresponding judgment threshold, and the judgment threshold of each quality assessment parameter is the threshold corresponding to each quality assessment parameter; for example, the threshold corresponding to the coefficient is 0.9 or 1.1, the threshold corresponding to the intercept is -0.1 or 0.1, the threshold corresponding to the determination coefficient is 0.6, and the threshold corresponding to the proportion of the absolute value of the difference in statistical data from different data platforms being less than 10% is 70%. It should be noted that the threshold corresponding to each quality assessment parameter can be adjusted according to actual conditions, and this application does not limit it specifically.
[0092] Specifically, the server compares multiple quality assessment parameters of the survey data of the two data platforms with the corresponding thresholds. If the multiple quality assessment parameters of the survey data of the two data platforms all meet the corresponding thresholds, it is determined that the quality identification results of the survey data of the two data platforms are the same quality level, indicating that the data quality of the survey data of the two data platforms meets the standards.
[0093] For example, in the linear equation y=ax+b, if the x coefficient a=0.9715, then the evaluation dimension meets the standard; if the intercept b=0.0131, then the evaluation dimension meets the standard; 2 =0.8374, indicating good regression results and thus meeting the criteria for this evaluation dimension. The proportion of absolute differences in statistical data from different data platforms less than 10% is 87.5%, indicating that this evaluation dimension meets the criteria. In this case, the survey data quality of both data platforms meets the criteria.
[0094] In this embodiment, the quality identification result of the survey data of the two-by-two data platforms is determined based on multiple quality evaluation parameters of the survey data of the two-by-two data platforms, which is conducive to improving the identification accuracy of the data quality.
[0095] In one embodiment, Figure 6 As shown, the above step S408, based on the quality identification results of the survey data of the two-to-two data platforms, determines the target quality identification results of the target survey data, which specifically includes the following steps:
[0096] Step S602: If the quality identification result shows that the quality levels are the same, the target quality identification result is determined to be up to standard.
[0097] Determining the target quality identification result as meeting the quality standard means that the target survey data meets the quality standard. Assuming there are two data platforms, if the quality identification results of the survey data on these two data platforms are the same quality level, then the target survey data is determined to meet the quality standard. Assuming there are more than two data platforms, if the quality identification results of the survey data on each of these data platforms are the same quality level, then the target survey data is determined to meet the quality standard.
[0098] For example, assuming there are three data platforms, namely data platform A, data platform B, and data platform C, and the quality identification results of the survey data of data platform A and data platform B are the same quality level, the quality identification results of the survey data of data platform A and data platform C are the same quality level, and the quality identification results of the survey data of data platform B and data platform C are the same quality level, then it means that the data quality of the target survey data composed of the survey data of data platform A, data platform B, and data platform C meets the standards.
[0099] Step S604: If the quality identification result is that the quality levels are different, the target quality identification result is determined as substandard quality.
[0100] Among them, determining the target quality identification result as substandard quality means that the data quality of the target survey data does not meet the standards; assuming that there are two data platforms, if the quality identification results of the survey data of these two data platforms are different quality levels, then the data quality of the target survey data is determined to be substandard; assuming that there are more than two data platforms, if among these data platforms, the quality identification results of the survey data of at least one group of two data platforms are different quality levels, then the data quality of the target survey data is determined to be substandard.
[0101] For example, suppose there are three data platforms, namely data platform A, data platform B and data platform C. The quality identification results of the survey data of data platform A and data platform B are the same quality level, the quality identification results of the survey data of data platform A and data platform C are different quality levels, and the quality identification results of the survey data of data platform B and data platform C are different quality levels. This means that the data quality of the target survey data composed of the survey data of data platform A, data platform B and data platform C does not meet the standards.
[0102] In this embodiment, the target quality identification result of the target survey data is determined based on the quality identification results of the survey data of each data platform, and the survey data collected by multiple data platforms are comprehensively considered, so that the quality identification of the survey data is more accurate, thereby improving the identification accuracy of the data quality.
[0103] In one embodiment, Figure 7 As shown, the above step S604, after determining that the target quality identification result is substandard, further includes the following steps:
[0104] Step S702: re-adjusting the data collection methods of at least two data platforms to obtain new data collection methods of at least two data platforms.
[0105] Step S704: Acquire survey data collected through at least two data platforms based on the new data collection method to obtain new target survey data.
[0106] Specifically, after the target quality identification result is determined to be substandard, the server re-filters the target survey data and, based on the filtered target survey data, determines the target quality identification result for the target survey data in accordance with steps S404-S410; and so on, when the number of times the target quality identification result is determined to be substandard reaches a preset number, such as 5 times, the data collection method of at least two data platforms is readjusted to obtain a new data collection method for at least two data platforms, so that at least two data platforms collect data based on the new data collection method, such as the questionnaire that was originally sent out in the form of a pop-up window is now sent out in the form of a push message; then the server obtains the survey data collected by at least two data platforms based on the new data collection method to obtain new target survey data.
[0107] In this embodiment, after the target quality identification result is determined to be substandard, the data collection methods of at least two data platforms are readjusted, and survey data collected by at least two data platforms based on the new data collection method is obtained to obtain new target survey data, which is conducive to ensuring that the data quality of the target survey data finally obtained meets the standards.
[0108] In one embodiment, after determining that the target quality identification result is up to standard, the above step S602 further includes the following steps: sending the target survey data to a corresponding data analysis platform; the data analysis platform is used to analyze the target survey data to obtain corresponding data analysis results.
[0109] For example, in the question "Which functions of W Pay have you used in the past year?", by analyzing the target survey data through the data analysis platform, we can obtain the user's usage of W Pay functions.
[0110] In this embodiment, the target survey data is sent to the corresponding data analysis platform only after the target quality identification result is determined to be up to standard, and the target survey data is analyzed by the data analysis platform, which is conducive to improving the accuracy of the data analysis result.
[0111] In one embodiment, Figure 8 As shown in the figure, another data quality identification method is provided, which is applied to Figure 3 The following steps are used as an example to illustrate the server in the example:
[0112] Step S802, obtaining target survey data; the target survey data includes survey data collected by at least two data platforms; the survey data of each data platform is collected by the corresponding data platform for the same user group based on the same questionnaire, including survey indicators and statistical data corresponding to the survey indicators.
[0113] Step S804: Filter invalid data in the target survey data to obtain filtered target survey data.
[0114] Step S806: Obtain a statistical data set of each survey indicator of the pairwise data platform based on the filtered target survey data.
[0115] Step S808: Perform linear regression processing on the statistical data sets of each survey indicator of each data platform to obtain a mapping relationship between the statistical data of each data platform.
[0116] Step S810: Based on the statistical data set and mapping relationship of each survey indicator of the two-by-two data platform, the regression fitting parameters of multiple preset dimensions of the survey data of the two-by-two data platform are obtained, and the regression fitting parameters of the multiple preset dimensions are used as multiple quality evaluation parameters of the survey data of the two-by-two data platform.
[0117] Step S812: If the multiple quality assessment parameters of the survey data of the two data platforms all meet the corresponding thresholds, it is determined that the quality identification results of the survey data of the two data platforms are the same quality level.
[0118] Step S814: If the quality identification result shows that the quality levels are the same, the target quality identification result is determined to be up to standard.
[0119] In the above-mentioned data quality identification method, the purpose of determining the target quality identification results of the target survey data based on the quality identification results of the survey data of each data platform is achieved. The survey data collected by multiple data platforms are comprehensively considered, making the quality identification of the survey data more accurate, thereby improving the recognition accuracy of data quality.
[0120] In one embodiment, the present application further provides a method for identifying data quality in a questionnaire survey scenario, which specifically includes the following contents:
[0121] (1) Questionnaire design: Design a corresponding questionnaire based on the information to be surveyed. For example, in the question “Which functions of W Pay have you used in the past year?”, the corresponding survey indicators are: transfer to friends, transfer to bank cards, transfer to mobile phone numbers, QR code payment, etc.
[0122] (2) Questionnaire delivery: Through online or offline interviews, users are asked to fill out questionnaires, so that users can see the questionnaires and fill in the corresponding questions; questionnaire delivery generally has certain sampling requirements, such as netizens, or female user groups aged 18-35, etc.
[0123] (3) Questionnaire collection: Generally, surveys will collect hundreds to thousands of user feedback samples based on specific needs, while the sample size for research with a better ROI is around 1,000.
[0124] (4) Questionnaire processing: Through certain technical means, such as judging the time required to fill in the questionnaire and logical checking of inconsistencies in the user's answers, some bad samples (such as invalid samples) are removed; through weight control, the unreasonable distribution of some key attributes of the recovered samples is adjusted.
[0125] (5) Data modeling: Through linear regression, the mapping relationship between channel A and channel B is established to obtain a linear equation (y = ax + b), where a is the x coefficient and b is the intercept.
[0126] (6) Quality assessment: Figure 9 , here there are 180 questions and 930 data points in total. We need to judge whether the scatter points are evenly distributed or not, and whether the linear regression fit is good or not. First, we calculate R 2 =0.8374, indicating that the regression effect is good and the evaluation dimension D1 meets the standard; the x coefficient a = 0.9715, the evaluation dimension D2 also meets the standard; the intercept b = 0.0131, the evaluation dimension D3 also meets the standard; intuitively, we can also see that channel A and channel B are evenly distributed on both sides of the y = x line, which is the expected distribution result. Figure 10 ,When looking at the difference between the two channels, we found that 87.5% of the points had an absolute ,difference value less than 10%, and the evaluation dimension D4 also met ,the standard.
[0127] (7) If the evaluation dimensions D1, D2, D3 and D4 all meet the standards, it means that the data quality of the survey data of Channel A and Channel B meets the standards, and then we can enter the stage of statistical analysis and data interpretation; otherwise, we need to return to the data processing stage, continue data processing and then conduct quality assessment. If after multiple iterations (for example, 5 times), it still fails to pass the quality assessment, it is necessary to return to the questionnaire delivery stage.
[0128] In this embodiment, by providing a quantifiable multi-channel data fusion technology solution for quality assessment, it is convenient for data analysts to judge and apply data more accurately, improve research quality and research efficiency, and generate greater data value (business value); of course, this method can also be applied to other scenarios, which are not specifically limited in this application.
[0129] It should be understood that although Figure 4-8 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 4-8 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The order of execution of these steps or stages is not necessarily one by one, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.
[0130] In one embodiment, Figure 11 As shown, a data quality identification device 1100 is provided. The device 1100 can be implemented as a software module or a hardware module, or a combination of both to form a part of a computer device. The device specifically includes: a data acquisition module 1102, a set acquisition module 1104, a parameter acquisition module 1106, a result determination module 1108, and a quality identification module 1110, wherein:
[0131] The data acquisition module 1102 is used to obtain target survey data; the target survey data includes survey data collected through at least two data platforms; the survey data of each data platform is collected by the corresponding data platform for the same user group based on the same questionnaire, including survey indicators and statistical data corresponding to the survey indicators.
[0132] The set acquisition module 1104 is used to obtain a statistical data set of each survey indicator of the pairwise data platform according to the target survey data.
[0133] The parameter acquisition module 1106 is used to perform linear regression processing on the statistical data set of each survey indicator of the two-by-two data platform to obtain multiple quality assessment parameters of the survey data of the two-by-two data platform.
[0134] The result determination module 1108 is configured to determine a quality identification result of the survey data of the two-by-two data platform according to a plurality of quality evaluation parameters of the survey data of the two-by-two data platform.
[0135] The quality identification module 1110 is configured to determine a target quality identification result for the target survey data based on the quality identification results for the survey data of the pairwise data platforms.
[0136] In one embodiment, the data quality identification device 1100 specifically further includes: a data filtering module;
[0137] A data filtering module is used to filter invalid data in the target survey data to obtain filtered target survey data;
[0138] The set acquisition module 1104 is further configured to acquire a statistical data set of each survey indicator of the pairwise data platform based on the filtered target survey data.
[0139] In one embodiment, the parameter acquisition module 1106 is also used to perform linear regression processing on the statistical data sets of each survey indicator of the two-by-two data platforms to obtain a mapping relationship between the statistical data of the two-by-two data platforms; based on the statistical data sets and mapping relationships of each survey indicator of the two-by-two data platforms, regression fitting parameters of multiple preset dimensions of the survey data of the two-by-two data platforms are obtained, and the regression fitting parameters of the multiple preset dimensions are used as multiple quality assessment parameters of the survey data of the two-by-two data platforms.
[0140] In one embodiment, the result determination module 1108 is further configured to determine that the quality identification results of the survey data of the two data platforms are of the same quality level if multiple quality assessment parameters of the survey data of the two data platforms all meet corresponding thresholds.
[0141] In one embodiment, the quality identification module 1110 is further configured to determine the target quality identification result as meeting the quality standard if the quality identification result shows that the quality levels are the same; and to determine the target quality identification result as failing to meet the quality standard if the quality identification result shows that the quality levels are different.
[0142] In one embodiment, the data quality identification device 1100 further includes: a re-collection module;
[0143] The re-collection module is used to readjust the data collection method of at least two data platforms to obtain a new data collection method for at least two data platforms; obtain the survey data collected by at least two data platforms based on the new data collection method to obtain new target survey data.
[0144] In one embodiment, the data quality identification device 1100 further includes: a data sending module;
[0145] The data sending module is used to send the target survey data to the corresponding data analysis platform; the data analysis platform is used to analyze the target survey data and obtain corresponding data analysis results.
[0146] For the specific definition of the data quality identification device, please refer to the definition of the data quality identification method above, which will not be repeated here. The various modules in the above-mentioned data quality identification device can be implemented in whole or in part by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0147] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 12 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data such as target survey data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a data quality identification method.
[0148] Those skilled in the art will understand that Figure 12 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0149] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0150] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.
[0151] In one embodiment, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of each of the above-described method embodiments.
[0152] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0153] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0154] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A data quality identification method, characterized in that: The method comprises: Obtaining target survey data; the target survey data includes survey data collected from at least two data platforms; the survey data of each data platform is collected from the same user group using the same questionnaire by the corresponding data platform, including survey indicators and statistical data corresponding to the survey indicators; According to the target survey data, a statistical data set of each survey indicator of the two-two data platform is obtained; Performing linear regression processing on the statistical data set of each survey indicator of the pairwise data platform to obtain multiple quality assessment parameters of the survey data of the pairwise data platform; Determining, based on a plurality of quality assessment parameters of the survey data of the pairwise data platforms, a quality identification result of the survey data of the pairwise data platforms, wherein the quality identification result is used to indicate whether the survey data of the pairwise data platforms have the same quality level; According to the quality identification results of the survey data of the pairwise data platforms, a target quality identification result of the target survey data is determined, and the target quality identification result is used to indicate whether the data quality of the target survey data meets the standards.
2. The method according to claim 1, characterized in that Before obtaining a statistical data set of each survey indicator of the two-by-two data platform based on the target survey data, the method further includes: filtering invalid data in the target survey data to obtain filtered target survey data; The step of obtaining a statistical data set of each survey indicator of the two-by-two data platform based on the target survey data includes: According to the filtered target survey data, a statistical data set of each survey indicator of the two-by-two data platform is obtained.
3. The method according to claim 1, characterized in that The linear regression processing is performed on the statistical data set of each survey indicator of the pairwise data platform to obtain multiple quality assessment parameters of the survey data of the pairwise data platform, including: Performing linear regression processing on the statistical data sets of each survey indicator of the two-by-two data platforms to obtain a mapping relationship between the statistical data of the two-by-two data platforms; Based on the statistical data set of each survey indicator of the two-by-two data platform and the mapping relationship, regression fitting parameters of multiple preset dimensions of the survey data of the two-by-two data platform are obtained, and the regression fitting parameters of the multiple preset dimensions are used as multiple quality assessment parameters of the survey data of the two-by-two data platform.
4. The method according to claim 1, wherein Determining a quality identification result of the survey data of the pairwise data platform according to a plurality of quality assessment parameters of the survey data of the pairwise data platform includes: If the multiple quality assessment parameters of the survey data of the two data platforms all meet the corresponding thresholds, it is determined that the quality identification results of the survey data of the two data platforms are the same quality level.
5. The method according to claim 1, wherein Determining a target quality identification result of the target survey data based on the quality identification result of the survey data of the pairwise data platform includes: If the quality identification result is that the quality level is the same, then the target quality identification result is determined to be up to standard; If the quality identification result is that the quality levels are different, the target quality identification result is determined to be substandard.
6. The method according to claim 5, characterized in that After determining that the target quality identification result is substandard, the method further includes: Re-adjusting the data collection methods of the at least two data platforms to obtain new data collection methods for the at least two data platforms; The survey data collected by the at least two data platforms based on the new data collection method is obtained to obtain new target survey data.
7. The method according to claim 5, characterized in that After determining that the target quality identification result is up to standard, the method further includes: The target survey data is sent to a corresponding data analysis platform; the data analysis platform is used to analyze the target survey data to obtain corresponding data analysis results.
8. A data quality identification device, characterized in that: The device comprises: A data acquisition module is configured to acquire target survey data; the target survey data includes survey data collected via at least two data platforms; the survey data on each data platform is collected from the same user group using the same questionnaire, including survey indicators and statistical data corresponding to the survey indicators; A set acquisition module is used to obtain a statistical data set of each survey indicator of the two-by-two data platform according to the target survey data; a parameter acquisition module, configured to perform linear regression processing on a statistical data set of each survey indicator of the pairwise data platform to obtain a plurality of quality assessment parameters of the survey data of the pairwise data platform; a result determination module, configured to determine a quality identification result of the survey data of the two-by-two data platforms based on a plurality of quality assessment parameters of the survey data of the two-by-two data platforms, wherein the quality identification result is used to indicate whether the survey data of the two-by-two data platforms have the same quality level; The quality identification module is used to determine the target quality identification result of the target survey data based on the quality identification result of the survey data of the pairwise data platform, and the target quality identification result is used to indicate whether the data quality of the target survey data meets the standard.
9. The data quality identification device according to claim 8, characterized in that: The device also includes a data filtering module, which is used to filter invalid data in the target survey data to obtain filtered target survey data; the set acquisition module is also used to obtain a statistical data set of each survey indicator of the pairwise data platform based on the filtered target survey data.
10. The data quality identification device according to claim 8, characterized in that: The parameter acquisition module is also used to perform linear regression processing on the statistical data sets of each survey indicator of the pairwise data platform to obtain a mapping relationship between the statistical data of the pairwise data platforms; based on the statistical data sets of each survey indicator of the pairwise data platform and the mapping relationship, the regression fitting parameters of multiple preset dimensions of the survey data of the pairwise data platform are obtained, and the regression fitting parameters of the multiple preset dimensions are used as multiple quality assessment parameters of the survey data of the pairwise data platform.
11. The data quality identification device according to claim 8, characterized in that: The result determination module is further configured to determine that the quality identification results of the survey data of the two data platforms are of the same quality level if multiple quality assessment parameters of the survey data of the two data platforms all meet corresponding thresholds.
12. The data quality identification device according to claim 8, characterized in that: The quality identification module is further configured to determine the target quality identification result as meeting the quality standard if the quality identification result shows that the quality levels are the same; and to determine the target quality identification result as failing to meet the quality standard if the quality identification result shows that the quality levels are different.
13. The data quality identification device according to claim 8, characterized in that: The device also includes a re-collection module, which is used to readjust the data collection method of the at least two data platforms to obtain a new data collection method for the at least two data platforms; obtain the survey data collected by the at least two data platforms based on the new data collection method to obtain new target survey data.
14. The data quality identification device according to claim 8, characterized in that: The device further includes a data sending module, which is used to send the target survey data to a corresponding data analysis platform; the data analysis platform is used to analyze the target survey data to obtain corresponding data analysis results.
15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
16. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
17. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Multi-channel analysis method and device based on analytic hierarchy process
CN103886168A
Object-oriented evaluation method and device
CN109978304A