Data verification method, data verification device, electronic device and storage medium

By using artificial intelligence technology to classify, extract features and predict data, the problem of low efficiency of manual verification is solved, and efficient and accurate data verification and correction are achieved.

CN115062705BActive Publication Date: 2025-09-19CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210689220.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-17
Publication Date
2025-09-19
Estimated Expiration
2042-06-17

AI Technical Summary

Technical Problem

In the existing technology, data verification mainly relies on manual methods, which leads to heavy workload and affects efficiency.

Method used

Artificial intelligence technology is used to obtain original business data, classify it using preset labels, extract key business features, and use preset algorithms to make predictions, generate business deviation data, and finally output verification results or perform data corrections.

Benefits of technology

It improves the efficiency and accuracy of data verification, shortens processing time, reduces calculation errors, enables timely measures to be taken, and improves the timeliness of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115062705B_ABST
    Figure CN115062705B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a data verification method, a data verification device, an electronic device, and a storage medium, which belong to the field of artificial intelligence technology. The method includes: obtaining the original business data of the target stage; classifying and processing the original business data according to preset labels to obtain labeled business data; extracting features from the labeled business data according to a preset extraction method to obtain business key features; predicting and processing the business key features through a preset algorithm to obtain business deviation data; performing status verification on the original business data according to the business deviation data to obtain verification results; outputting task scheduling instructions or performing data correction on the original business data according to the verification results. The embodiments of the present application can improve data verification efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data verification method, a data verification device, an electronic device, and a storage medium. Background Art

[0002] At present, manual verification is often used when verifying data. This method often requires manual verification of the data at each stage one by one, which is a large workload and affects the efficiency of data verification. Therefore, how to improve the efficiency of data verification has become a technical problem that needs to be solved urgently. Summary of the Invention

[0003] The main purpose of the embodiments of the present application is to propose a data verification method, a data verification device, an electronic device and a storage medium, aiming to improve the efficiency of data verification.

[0004] To achieve the above objectives, a first aspect of an embodiment of the present application provides a data verification method, the method comprising:

[0005] Obtain the original business data of the target stage;

[0006] Classify the original business data according to preset labels to obtain labeled business data;

[0007] Extracting features from the tag business data according to a preset extraction method to obtain key business features;

[0008] Predicting and processing the key business features through a preset algorithm to obtain business deviation data;

[0009] Performing a status check on the original business data according to the business deviation data to obtain a check result;

[0010] Output a task scheduling instruction or perform data correction on the original business data according to the verification result.

[0011] In some embodiments, the step of classifying the original business data according to preset tags to obtain tagged business data includes:

[0012] Acquiring attribute data of the target stage;

[0013] Filter the preset tags in the preset tag library according to the attribute data to obtain the target tag;

[0014] Performing label probability calculation on the original service data using a preset function and the target label to obtain a label probability value;

[0015] The original service data is classified according to the label probability value to obtain the label service data.

[0016] In some embodiments, the target stage includes a first stage, the extraction method includes a first method, and the step of extracting features from the tagged business data according to a preset extraction method to obtain key business features includes:

[0017] If the label business data is the business data of the first phase, obtaining the data category segment corresponding to the first method;

[0018] Extract keywords from the tagged business data according to the data category segment to obtain initial business keywords;

[0019] The initial business keywords are filtered to obtain the business key features.

[0020] In some embodiments, the target stage includes a second stage, the extraction method includes a second method, and the step of extracting features from the tagged business data according to a preset extraction method to obtain key business features includes:

[0021] If the label business data is business data of the second phase, then according to the preset data dimension, an algorithm configuration table corresponding to the second method is obtained, where the algorithm configuration table stores preset labels and preset algorithms, and there is a mapping relationship between the preset labels and the preset algorithms;

[0022] Extracting a target tag corresponding to the tag service data from the preset tag;

[0023] According to the mapping relationship, extracting the preset algorithm corresponding to the target tag from the algorithm configuration table;

[0024] Perform field extraction on the tag service data using the preset algorithm to obtain initial key fields;

[0025] Calculate the field value of the initial key field to obtain the business key feature.

[0026] In some embodiments, the step of performing prediction processing on the key business features using a preset algorithm to obtain business deviation data includes:

[0027] Perform regression prediction on the key business features using the least squares method to obtain initial prediction data;

[0028] The initial prediction data is calibrated according to a preset deviation parameter to obtain the business deviation data.

[0029] In some embodiments, the verification result includes a first result and a second result, and the step of performing status verification on the original business data according to the business deviation data to obtain the verification result includes:

[0030] If the business deviation data is within the preset deviation range, determining that the original business data is in a normal state and outputting a first result;

[0031] If the business deviation data is outside the preset deviation range, it is determined that the original business data is in an abnormal state and a second result is output.

[0032] In some embodiments, the verification result includes a first result and a second result, and the step of outputting a task scheduling instruction or performing data correction on the original business data according to the verification result includes:

[0033] If the verification result is the first result, outputting a task scheduling instruction to verify the business data according to the task scheduling instruction;

[0034] If the verification result is the second result, the original business data is anomaly located, anomaly information is generated, and data correction is performed on the original business data based on the anomaly information.

[0035] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a data verification device, comprising:

[0036] Data acquisition module, used to obtain the original business data of the target stage;

[0037] A classification module, configured to classify the original business data according to preset labels to obtain labeled business data;

[0038] A feature extraction module is used to extract features from the tag business data according to a preset extraction method to obtain key business features;

[0039] A prediction module is used to predict the key business features using a preset algorithm to obtain business deviation data;

[0040] A verification module, configured to perform a status verification on the original business data according to the business deviation data to obtain a verification result;

[0041] A processing module is used to output a task scheduling instruction or perform data correction on the original business data according to the verification result.

[0042] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory, a processor, a program stored on the memory and runnable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the method described in the first aspect above is implemented.

[0043] To achieve the above-mentioned purpose, the fourth aspect of the embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement the steps of the data verification method of the first aspect.

[0044] The data verification method, data verification device, electronic device and storage medium proposed in the present application obtain the original business data of the target stage; classify and process the original business data according to preset labels to obtain labeled business data, and can obtain business data of different target stages according to business needs, thereby performing data verification at different business stages and improving the data accuracy of the entire process. Furthermore, according to a preset extraction method, feature extraction is performed on the labeled business data to obtain business key features; the business key features are predicted and processed by a preset algorithm to obtain business deviation data, and different extraction methods can be adopted for business data of different target stages, so that the acquired business key features are more accurate. Compared with the manual verification method, the embodiment of the present application performs data prediction through a preset algorithm, which can effectively shorten the data processing time and reduce the calculation error of business deviation data, thereby improving the efficiency and accuracy of data verification. Finally, the status of the original business data is checked based on the business deviation data to obtain the verification results; based on the verification results, task scheduling instructions are output or data correction is performed on the original business data. Based on the business deviation data, it can be analyzed whether the original business data is in a normal state or an abnormal state, so that corresponding measures can be taken in a timely manner according to the verification results, thereby improving the timeliness of data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flow chart of the data verification method provided in an embodiment of the present application;

[0046] Figure 2 yes Figure 1 Flowchart of step S102 in FIG.

[0047] Figure 3 yes Figure 1 Flowchart of step S103 in FIG.

[0048] Figure 4 yes Figure 1Another flowchart of step S103 in FIG.

[0049] Figure 5 yes Figure 1 Flowchart of step S104 in FIG.

[0050] Figure 6 yes Figure 1 Flowchart of step S105 in FIG.

[0051] Figure 7 yes Figure 1 Flowchart of step S106 in FIG.

[0052] Figure 8 Schematic diagram of the structure of the data verification device provided in the embodiment of the present application;

[0053] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0055] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0057] First, let’s analyze some of the terms used in this application:

[0058] Artificial intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It also encompasses the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.

[0059] Natural language processing (NLP): NLP uses computers to process, understand, and apply human languages ​​(such as Chinese and English). A branch of artificial intelligence, NLP is an interdisciplinary field between computer science and linguistics, often referred to as computational linguistics. Natural language processing encompasses grammatical analysis, semantic analysis, and discourse comprehension. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It encompasses data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistics research related to language computing.

[0060] Information Extraction: A text processing technology that extracts specified types of entity, relationship, event, and other factual information from natural language text and forms structured data output. Information extraction is a technology that extracts specific information from text data. Text data is composed of some specific units, such as sentences, paragraphs, and chapters. Text information is composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or a combination of these specific units. Extracting noun phrases, names, place names, etc. from text data is all text information extraction. Of course, the information extracted by text information extraction technology can be of various types.

[0061] MP (Mass-Production): mass production.

[0062] Least squares (generalized least squares): Also known as the method of least squares, this is a mathematical optimization technique. It seeks the best function matching data by minimizing the sum of squared errors. Least squares can be used to easily find unknown data and minimize the sum of squared errors between the found data and the actual data. Least squares can also be used for curve fitting. Other optimization problems can also be formulated using least squares methods, such as minimizing energy or maximizing entropy.

[0063] Data fitting, also known as curve fitting or simply curve fitting, is a mathematical method that substitutes existing data into a numerical expression. Scientific and engineering problems can be solved by obtaining discrete data through methods such as sampling and experimentation. Based on this data, we often hope to find a continuous function (i.e., a curve) or a more complex discrete equation that matches the known data. This process is called fitting.

[0064] Currently, manual verification is often used to verify data in insurance business scenarios. This method often requires manual verification of data at each stage, which is a large workload and affects the efficiency of data verification. Therefore, how to improve the efficiency of data verification has become a technical problem that needs to be solved urgently.

[0065] Based on this, the embodiments of the present application provide a data verification method, a data verification device, an electronic device and a storage medium, aiming to improve the efficiency of data verification.

[0066] The data verification method, data verification device, electronic device and storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the data verification method in the embodiments of the present application is described.

[0067] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0068] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0069] The data verification method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The data verification method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the data verification method, etc., but is not limited to the above forms.

[0070] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in computer-readable storage media of local and remote computers, including storage devices.

[0071] Figure 1 This is an optional flow chart of the data verification method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S106.

[0072] Step S101, obtaining the original business data of the target stage;

[0073] Step S102: classify the original business data according to the preset tags to obtain labeled business data;

[0074] Step S103: extracting features from the tag business data according to a preset extraction method to obtain key business features;

[0075] Step S104: predicting the key business features using a preset algorithm to obtain business deviation data;

[0076] Step S105: Check the status of the original business data based on the business deviation data to obtain a check result;

[0077] Step S106: outputting a task scheduling instruction or performing data correction on the original business data according to the verification result.

[0078] Steps S101 to S106 shown in the embodiment of the present application, by obtaining the original business data of the target stage; classifying and processing the original business data according to preset labels to obtain labeled business data, can obtain business data of different target stages according to business needs, thereby performing data verification at different business stages and improving the data accuracy of the entire process. Furthermore, according to the preset extraction method, feature extraction is performed on the labeled business data to obtain business key features; the business key features are predicted and processed by the preset algorithm to obtain business deviation data, and different extraction methods can be adopted for the business data of different target stages, so that the acquired business key features are more accurate. Compared with the manual verification method, the embodiment of the present application uses a preset algorithm to perform data prediction, which can effectively shorten the data processing time and reduce the calculation error of the business deviation data, thereby improving the efficiency and accuracy of data verification. Finally, the status of the original business data is checked based on the business deviation data to obtain the verification results; based on the verification results, task scheduling instructions are output or data correction is performed on the original business data. Based on the business deviation data, it can be analyzed whether the original business data is in a normal state or an abnormal state, so that corresponding measures can be taken in a timely manner according to the verification results, thereby improving the timeliness of data processing.

[0079] In step S101 of some embodiments, a web crawler can be written and, after setting up a data source, targeted data can be crawled to obtain raw business data for the target stage. Raw business data can also be obtained through other methods, but is not limited to these. It should be noted that in the insurance business, the target stages primarily include the policy aggregation stage, the calculation stage, and the data compression stage, each of which generates corresponding business result data. For example, in the policy aggregation stage, the raw business data includes multiple types of insurance data, including the number of policies, premiums, and so on.

[0080] See also Figure 2 In some embodiments, step S102 may include but is not limited to steps S201 to S204:

[0081] Step S201, obtaining attribute data of the target stage;

[0082] Step S202: Filter the preset tags in the preset tag library according to the attribute data to obtain a target tag;

[0083] Step S203, performing label probability calculation on the original business data using a preset function and the target label to obtain a label probability value;

[0084] Step S204 : classify the original service data according to the label probability value to obtain label service data.

[0085] In step S201 of some embodiments, the target phase includes a policy aggregation phase, a calculation phase, and a data compression phase. Each phase generates a large amount of business data. Furthermore, because the attribute data for each phase differs, different attribute data can be used to distinguish between different phases. For example, when the target phase is the policy aggregation phase, the attribute data includes, for example, policy type; when the target phase is the data compression phase, the attribute data includes, for example, product type. Therefore, the attribute data for the target phase can be obtained by programming a web crawler, setting up a data source, and then performing targeted data crawling. Attribute data can also be obtained through other methods, without limitation.

[0086] In step S202 of some embodiments, the preset tag library stores data tags for different target stages, i.e., preset tags, where the preset tags include insurance category tags, time tags, quantity tags, or numerical tags, etc. Using the attribute data of different stages, preset tags that meet the requirements of the current stage can be filtered out from multiple preset tags, and these preset tags can be used as target tags. For example, in the policy aggregation stage, the attribute data includes the policy type. Based on the policy type, the preset tags can be filtered out to select quantity tags and numerical tags, so that the quantity of policies of different types can be obtained based on the quantity tags, and the insurance amounts of different types of policies can be obtained based on the numerical tags.

[0087] In step S203 of some embodiments, the preset function is a softmax function, which can conveniently create a probability distribution of the original business data on the target label, thereby obtaining a label probability value for each target label. The label probability value can be used to reflect the possibility that the original business data belongs to the target label.

[0088] In step S204 of some embodiments, since the label probability value can be used to reflect the possibility that the original business data belongs to the target label, the original business data is classified according to the label probability value, and the original business data is divided into the category of the target label with the largest label probability value, so as to distinguish the original business data at different stages and obtain different types of labeled business data at different stages. For example, in this way, the original business data belonging to different types of insurance summarized at the policy summary stage can be more conveniently divided according to the insurance category to obtain the labeled business data belonging to the policy summary stage.

[0089] See also Figure 3 In some embodiments, the target stage includes the first stage, the extraction method includes the first method, and step S103 may include but is not limited to steps S301 to S303:

[0090] Step S301: If the tagged business data is business data of the first phase, then obtain the data category segment corresponding to the first method;

[0091] Step S302: extract keywords from the tagged business data according to the data category segment to obtain initial business keywords;

[0092] Step S303: Filter the initial business keywords to obtain key business features.

[0093] In step S301 of some embodiments, the first stage is the policy summary stage. If the labeled business data is business data in the policy summary stage, the data category segment corresponding to the first method is obtained. The first method is generally to set corresponding label segments for different types of data, that is, data type segments, where the data type segments are specifically the number of policies, policy amount, etc.

[0094] In step S302 of some embodiments, the label business data can be processed into several sentence nodes by the TF-IDF algorithm. Specifically, the frequency of occurrence of the data category segment in each sentence in the label business data is calculated by the TF-IDF algorithm to obtain the term frequency (TF) of the data category segment, where TF = the number of times the data category segment appears / the number of segments in the label business data; further, the inverse document frequency (IDF) of the data category segment is calculated, where IDF = log (the total number of segments in the label business data / (the number of segments in the label business data containing the data category segment + 1)). Finally, the comprehensive frequency value of the segment is calculated based on the term frequency and the inverse document frequency, where the comprehensive frequency value = term frequency * inverse document frequency. The segment with the largest comprehensive frequency value is selected in each target text as the initial business keyword.

[0095] In step S303 of some embodiments, other word segments within the preset character range of the initial business keyword are extracted, and the word segments with higher relevance to the initial business keyword are retained, and the word segments with lower relevance are filtered out to obtain business key features. For example, if the initial business keyword is the policy amount, the word segments close to the policy amount field position are extracted, and the word segments with numerical values ​​and amount units among the extracted word segments are retained to obtain the business key features corresponding to the policy amount: 50,000 yuan.

[0096] See also Figure 4In some embodiments, the target stage includes the second stage, the extraction method includes the second method, and step S103 may include but is not limited to steps S401 to S405:

[0097] Step S401: If the tagged business data is business data of the second phase, then according to the preset data dimension, an algorithm configuration table corresponding to the second method is obtained, where the algorithm configuration table stores preset tags and preset algorithms, and there is a mapping relationship between the preset tags and the preset algorithms;

[0098] Step S402: extracting a target tag corresponding to the tag service data from the preset tags;

[0099] Step S403: extracting the preset algorithm corresponding to the target tag from the algorithm configuration table according to the mapping relationship;

[0100] Step S404: extracting fields from the tag service data using a preset algorithm to obtain initial key fields;

[0101] Step S405: Calculate the field values ​​of the initial key fields to obtain business key features.

[0102] In step S401 of some embodiments, the second stage may refer to the calculation stage or the data compression stage. The second method is generally an algorithm configuration method. The data dimension can be set according to actual business needs. When in the calculation stage, the preset data dimension is the data analysis at the policy level. When in the data compression stage, the preset data dimension is the data analysis at the product level. When the label business is actually the business data in the calculation stage or the business data in the data compression stage. Obtain the algorithm configuration table corresponding to the second method, wherein the algorithm configuration table stores preset labels and preset algorithms, and there is a mapping relationship between the preset labels and the preset algorithms. It should be noted that this mapping relationship can be in the form of one-to-one or one-to-many.

[0103] In step S402 of some embodiments, since the target tag is one of the preset tags, the same tag can be conveniently extracted from the preset tags according to the tag name or tag value of the target tag.

[0104] In some embodiments, in step S403, after determining which of the preset tags is the target tag, the preset algorithm corresponding to the preset tag and the mapping relationship data between the preset algorithm and the preset algorithm can be directly extracted from the algorithm configuration table. For example, the algorithm configuration table is traversed based on the target tag to extract the preset algorithms corresponding to the target tag and the mapping data between the target tag and the preset algorithms, where the mapping data includes the target tag, the preset algorithm, and the correlation between the target tag and the preset algorithm.

[0105] In step S404 of some embodiments, the preset algorithm may be to extract the fields of the tag business data through the slice function, the start function, and the stop function to obtain the initial key fields, wherein the preset algorithm may be the TF-IDF algorithm, etc., without limitation.

[0106] In step S405 of some embodiments, when calculating the field value of the initial key field, the field value of the initial key field is first extracted by the tan function, and then the field values ​​of all calculated initial key fields are summed or averaged to obtain the business key feature.

[0107] In one specific embodiment, when data verification is in the calculation phase, a preset algorithm configuration table is retrieved. This table configures the target tag and the corresponding preset algorithm. This table specifies the fields to be calculated and the method to be used to calculate them. During the actual calculation of key business features, field extraction and field value calculation can be performed on the tagged business data according to the corresponding preset algorithm to obtain the key business features.

[0108] See also Figure 5 In some embodiments, step S104 may include but is not limited to steps S501 to S502:

[0109] Step S501, performing regression prediction on the key business features by the least squares method to obtain initial prediction data;

[0110] Step S502: calibrate the initial prediction data according to the preset deviation parameters to obtain business deviation data.

[0111] In step S501 of some embodiments, the least squares method can be used to minimize the sum of squares of errors to find the best function match for the data. Therefore, the least squares method can be used to perform regression prediction on key business features, generate a regression line, and obtain initial prediction data, which can be expressed in the form of a curve.

[0112] In step S502 of some embodiments, the preset deviation parameter can be set according to actual business needs without restriction. For example, if the deviation parameter is ±5%, the regression prediction range corresponding to the initial prediction data is calibrated using the preset deviation parameter to obtain business deviation data, which is the regression prediction value range.

[0113] See also Figure 6 In some embodiments, the verification result includes a first result and a second result, and step S105 includes but is not limited to steps S601 to S602:

[0114] Step S601: If the business deviation data is within a preset deviation range, the original business data is determined to be in a normal state, and a first result is output;

[0115] Step S602: If the business deviation data is outside the preset deviation range, the original business data is determined to be in an abnormal state, and a second result is output.

[0116] In some embodiments, in step S601, if the business deviation data is within a preset deviation range, that is, if the deviation of the field information corresponding to the tagged business data is within the preset deviation range, then the original business data is normal, and a first result is output. The specific semantic content of the first result includes: the current original business data is normal and the verification has passed. Thus, the first result can conveniently represent normal original business data.

[0117] In some embodiments, in step S602, if the business deviation data is outside a preset deviation range, that is, the deviation of the field information corresponding to the tagged business data is outside the preset deviation range, then the original business data is abnormal, and a second result is output. The specific semantic content of the second result includes: the current original business data is in an abnormal state, the verification fails, and the current original business data needs to be analyzed for abnormality, the cause of the abnormality is determined, and data correction operations are performed on the current original business data. Thus, the second result can conveniently characterize the abnormal original business data.

[0118] See also Figure 7 In some embodiments, the verification result includes a first result and a second result. Step S106 may include, but is not limited to, steps S701 to S702:

[0119] Step S701: If the verification result is the first result, output a task scheduling instruction to verify the business data according to the task scheduling instruction;

[0120] Step S702: If the verification result is the second result, the original business data is anomaly located, anomaly information is generated, and data correction is performed on the original business data according to the anomaly information.

[0121] In step S701 of some embodiments, if the verification result is the first result, it means that the original business data of the current stage is normal and the data verification of the next stage can be carried out. Therefore, the task scheduling instruction is output, and the original business data of the next stage is extracted according to the task scheduling instruction, and the original business data of the next stage is verified with the same verification process.

[0122] In step S702 of some embodiments, if the verification result is the second result, it means that the original business data at the current stage is abnormal and the abnormal data needs to be processed. Therefore, based on the comparison between the original business data and the historical business data, the location of the original business data in the abnormal state is determined, and abnormal information is generated. The abnormal information includes the abnormal location, deviation, and abnormal cause, etc., so that the original business data is corrected according to the preset solution. The preset solution can be formulated based on the abnormal cause and solution determined according to the abnormal problems checked historically. The solution can be stored in the form of a configuration table, so that after determining the original business data in the abnormal state, the configuration table can be traversed to extract the corresponding solution.

[0123] In a specific embodiment, the process of verifying insurance business data for a certain month includes three stages, namely, the insurance policy aggregation stage, the DCS calculation stage, and the data compression stage.

[0124] First, we obtain raw business data from the policy aggregation phase, typically presented as a result table. We then categorize and count this data based on the insurance type attribute to obtain labeled business data for each insurance type. Furthermore, we extract key fields from this labeled business data based on data type terms such as policy quantity and policy amount to obtain key business features. These key features primarily include the specific number of policies and policy amounts for that month. Furthermore, these key business features are regressed and predicted by the least squares method, and a regression curve corresponding to each key business feature is constructed. A deviation of ±5% is assigned to each key business feature to obtain the business deviation data of each key business feature. The business deviation data is compared with the key business feature of that month in the historical records. If it exceeds the preset deviation range, the original business data corresponding to the key business feature is determined to be in an abnormal state, and an abnormal analysis is performed on this original business data. Specifically, the preset problem configuration table is traversed according to the key business features, and the abnormal causes and solutions determined based on the abnormal problems checked historically are retrieved from the problem configuration table, so as to correct the original business data in the abnormal state until all the original business data in the policy summary stage are in a normal state, output the task scheduling instructions, and extract the original business data of the next stage according to the task scheduling instructions.

[0125] Furthermore, after the verification of the original business data in the policy aggregation phase is completed, the original business data in the DCS calculation phase is verified and processed according to the task scheduling instructions. Specifically, the label business data of the DCS calculation phase is obtained, and the preset algorithm configuration table is obtained. The algorithm configuration table is configured with the target label and the preset algorithm corresponding to the target label, that is, the label business data to be calculated and the calculation method for calculating the label business data can be clearly defined through the algorithm configuration table. The preset algorithm corresponding to each label business data is extracted through the algorithm configuration table, and the field extraction and field value calculation of the label business data are performed according to the corresponding preset algorithm, thereby obtaining the business key features, wherein the business key features. Then, regression prediction is performed on these key business features through the least squares method, and a regression curve corresponding to each key business feature is constructed. A deviation of ±5% is assigned to each key business feature to obtain the business deviation data of each key business feature. The business deviation data is compared with the key business feature of that month in the historical records. If it exceeds the preset deviation range, the original business data corresponding to the key business feature is determined to be in an abnormal state, and an abnormal analysis is performed on this original business data. Specifically, the preset problem configuration table is traversed according to the key business features, and the abnormal causes and solutions determined based on the abnormal problems checked historically are retrieved from the problem configuration table, so as to correct the original business data in the abnormal state until all the original business data in the policy summary stage are in normal state, output the task scheduling instructions, and extract the original business data of the next stage according to the task scheduling instructions.

[0126] Furthermore, after the original business data of the DSC calculation phase is verified, the original business data of the data compression phase is verified and processed according to the task scheduling instructions. Specifically, the data verification in the data compression phase is mainly to verify the same type of data. First, a group by operation is performed on the same type of tagged business data according to the preset dimension to realize the aggregation of the tagged business data and obtain the field values ​​corresponding to the tagged business data. These field data are then summed or averaged to obtain the key business features of this phase. Similarly, the least squares method is used to perform abnormal verification on the key business features of this phase, so as to correct the original business data in an abnormal state.

[0127] It should be noted that in the embodiment of the present application, if the original business data of the current stage has not been verified, the original business data of the next stage will not be verified. When the original business data of the current stage is verified, it can be automatically verified by a computer program and enter the data verification of the next stage. In this way, the original business data of each stage can be checked one by one, which can effectively save the time of business personnel in verifying the result data and improve the efficiency of data verification. At the same time, each stage can be automatically verified by a computer program, and when the verification is qualified, it can automatically enter the data verification of the next stage, realizing seamless connection of data verification of each stage. In addition, when the original business data in an abnormal state is found, the cause of the possible problem can be queried from the preset problem configuration table by combining parameters with historical problems through a computer program, thereby facilitating business personnel to quickly locate abnormal problems and improve the timeliness of problem solving.

[0128] The data verification method of the embodiment of the present application obtains the original business data of the target stage; classifies and processes the original business data according to preset labels to obtain labeled business data. It can obtain business data of different target stages according to business needs, thereby performing data verification at different business stages and improving the data accuracy of the entire process. Furthermore, the labeled business data is feature extracted according to a preset extraction method to obtain business key features; the business key features are predicted and processed by a preset algorithm to obtain business deviation data. Different extraction methods can be used to extract business data of different target stages, thereby making the obtained business key features more accurate. Compared with the manual verification method, the embodiment of the present application uses a preset algorithm to predict data, which can effectively shorten the data processing time and reduce the calculation error of business deviation data, thereby improving the efficiency and accuracy of data verification. Finally, the status of the original business data is verified based on the business deviation data to obtain a verification result; based on the verification result, a task scheduling instruction is output or data correction is performed on the original business data. Based on the business deviation data, it can analyze whether the original business data is in a normal state or an abnormal state, so that corresponding measures can be taken in a timely manner according to the verification result, thereby improving the timeliness of data processing.

[0129] See also Figure 8 The present invention also provides a data verification device that can implement the above-mentioned data verification method. The device includes:

[0130] Data acquisition module 801, used to obtain the original business data of the target stage;

[0131] The classification module 802 is used to classify the original business data according to the preset labels to obtain labeled business data;

[0132] The feature extraction module 803 is used to extract features from the tag business data according to a preset extraction method to obtain key business features;

[0133] Prediction module 804, used to predict key business features using a preset algorithm to obtain business deviation data;

[0134] Verification module 805, used to verify the status of original business data based on the business deviation data and obtain verification results;

[0135] The processing module 806 is used to output a task scheduling instruction or perform data correction on the original business data according to the verification result.

[0136] In some embodiments, the classification module 802 includes:

[0137] An attribute data acquisition unit, used to acquire attribute data of a target stage;

[0138] A screening unit, configured to screen the preset tags in the preset tag library according to the attribute data to obtain a target tag;

[0139] The probability calculation unit is used to calculate the label probability of the original business data using a preset function and the target label to obtain a label probability value;

[0140] The data classification unit is used to classify the original business data according to the label probability value to obtain labeled business data.

[0141] In some embodiments, the target stage includes the first stage, the extraction method includes the first method, and the feature extraction module 803 includes:

[0142] a word segment obtaining unit, configured to obtain a data category word segment corresponding to a first method if the label business data is business data of the first phase;

[0143] A keyword extraction unit is used to extract keywords from the tagged business data according to the data category segment to obtain initial business keywords;

[0144] The filtering unit is used to filter the initial business keywords to obtain business key features.

[0145] In some embodiments, the target stage includes the second stage, the extraction method includes the second method, and the feature extraction module 803 further includes:

[0146] A configuration table acquisition unit is configured to acquire, if the tag business data is business data of the second phase, an algorithm configuration table corresponding to the second method according to a preset data dimension, wherein the algorithm configuration table stores preset tags and preset algorithms, and a mapping relationship exists between the preset tags and the preset algorithms;

[0147] A tag extraction unit, configured to extract a target tag corresponding to the tag service data from a preset tag;

[0148] An algorithm extraction unit is used to extract the preset algorithm corresponding to the target tag from the algorithm configuration table according to the mapping relationship;

[0149] A field extraction unit is used to extract fields from the tag business data using a preset algorithm to obtain initial key fields;

[0150] The field value calculation unit is used to calculate the field value of the initial key field to obtain the business key features.

[0151] In some embodiments, the prediction module 804 includes:

[0152] The regression prediction unit is used to perform regression prediction on key business features using the least squares method to obtain initial prediction data;

[0153] The calibration unit is used to calibrate the initial prediction data according to the preset deviation parameters to obtain business deviation data.

[0154] In some embodiments, the verification result includes a first result and a second result, and the verification module 805 includes:

[0155] A first judgment unit is configured to determine that the original business data is in a normal state if the business deviation data is within a preset deviation range, and output a first result;

[0156] The second judgment unit is configured to determine that the original business data is in an abnormal state if the business deviation data is outside a preset deviation range, and output a second result.

[0157] In some embodiments, the verification result includes a first result and a second result, and the processing module 806 includes:

[0158] an output unit, configured to output a task scheduling instruction if the verification result is the first result, so as to verify the business data according to the task scheduling instruction;

[0159] The exception processing unit is used to locate the exception of the original business data, generate exception information, and perform data correction on the original business data according to the exception information if the verification result is the second result.

[0160] The specific implementation of the data verification device is basically the same as the specific embodiment of the above-mentioned data verification method, and will not be repeated here.

[0161] The present application also provides an electronic device comprising: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for enabling communication between the processor and the memory, wherein the program, when executed by the processor, implements the aforementioned data verification method. The electronic device may be any intelligent terminal, such as a tablet computer or an in-vehicle computer.

[0162] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0163] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0164] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant computer program code is stored in the memory 902 and is called by the processor 901 to execute the data verification method of the embodiments of this application;

[0165] Input / output interface 903, used to implement information input and output;

[0166] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0167] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );

[0168] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0169] An embodiment of the present application also provides a computer-readable storage medium, which stores one or more computer programs. The one or more computer programs can be executed by one or more processors to implement the above-mentioned data verification method.

[0170] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0171] The data verification method, data verification device, electronic device and storage medium provided in the embodiment of the present application obtain the original business data of the target stage; classify and process the original business data according to preset labels to obtain labeled business data, and can obtain business data of different target stages according to business needs, thereby performing data verification at different business stages and improving the data accuracy of the entire process. Furthermore, according to the preset extraction method, feature extraction is performed on the labeled business data to obtain business key features; the business key features are predicted and processed by the preset algorithm to obtain business deviation data, and different extraction methods can be adopted for the business data of different target stages, so that the acquired business key features are more accurate. Compared with the manual verification method, the embodiment of the present application performs data prediction through a preset algorithm, which can effectively shorten the data processing time and reduce the calculation error of the business deviation data, thereby improving the efficiency and accuracy of data verification. Finally, the status of the original business data is checked based on the business deviation data to obtain the verification results; based on the verification results, task scheduling instructions are output or data correction is performed on the original business data. Based on the business deviation data, it can be analyzed whether the original business data is in a normal state or an abnormal state, so that corresponding measures can be taken in a timely manner according to the verification results, thereby improving the timeliness of data processing.

[0172] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0173] It will be understood by those skilled in the art that Figure 1-7The technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or a combination of certain steps, or different steps.

[0174] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0175] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0176] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0177] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0178] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0179] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0180] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0181] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a computer-readable storage medium, including multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned computer-readable storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store computer programs.

[0182] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A data verification method, characterized in that: The method comprises: Obtain the original business data of the target stage; Classify the original business data according to preset labels to obtain labeled business data; Extracting features from the tag business data according to a preset extraction method to obtain key business features; predicting and processing the key business features using a preset algorithm to obtain business deviation data; Performing a status check on the original business data according to the business deviation data to obtain a check result; outputting a task scheduling instruction or performing data correction on the original business data according to the check result; The target stage includes a first stage and a second stage, the extraction method includes a first method and a second method, the first stage is a policy aggregation stage, and the second stage is a calculation stage or a data compression stage; the feature extraction of the label business data according to the preset extraction method to obtain the key business features includes: If the tagged business data is business data of the first phase, then obtaining the data category word segment corresponding to the first method; performing keyword extraction on the tagged business data using the TF-IDF algorithm based on the data category word segment to obtain an initial business keyword; extracting other word segments within a preset character range of the initial business keyword, retaining word segments whose relevance to the initial business keyword is greater than a threshold, and obtaining the business key feature; If the tagged business data is the business data of the second phase, then according to the preset data dimension, the algorithm configuration table corresponding to the second method is obtained, the algorithm configuration table stores preset tags and preset algorithms, and there is a mapping relationship between the preset tags and the preset algorithms; according to the tag name or tag value of the target tag, the target tag corresponding to the tagged business data is extracted from the preset tag; according to the mapping relationship, the preset algorithm corresponding to the target tag is extracted from the algorithm configuration table; the field of the tagged business data is extracted by the preset algorithm to obtain the initial key field; the field value of the initial key field is extracted by the tan function, and the field values ​​of all calculated initial key fields are summed or averaged to obtain the business key feature; The predicting process of the key business features by a preset algorithm to obtain business deviation data includes: The business key features are regressed and predicted by the least squares method to generate a regression line and obtain initial prediction data; the regression prediction range corresponding to the initial prediction data is calibrated according to the preset deviation parameters to obtain the business deviation data, which is the regression prediction value interval.

2. The data verification method according to claim 1, characterized in that: The step of classifying the original business data according to the preset labels to obtain labeled business data includes: Acquiring attribute data of the target stage; Filter the preset tags in the preset tag library according to the attribute data to obtain the target tag; Performing label probability calculation on the original service data using a preset function and the target label to obtain a label probability value; The original service data is classified according to the label probability value to obtain the label service data.

3. The data verification method according to claim 1, characterized in that: The verification result includes a first result and a second result. The step of performing status verification on the original business data according to the business deviation data to obtain the verification result includes: If the business deviation data is within the preset deviation range, determining that the original business data is in a normal state and outputting a first result; If the business deviation data is outside the preset deviation range, it is determined that the original business data is in an abnormal state and a second result is output.

4. The data verification method according to any one of claims 1 to 3, characterized in that: The verification result includes a first result and a second result, and the step of outputting a task scheduling instruction or performing data correction on the original business data according to the verification result includes: If the verification result is the first result, outputting a task scheduling instruction to verify the business data according to the task scheduling instruction; If the verification result is the second result, the original business data is anomaly located, anomaly information is generated, and data correction is performed on the original business data based on the anomaly information.

5. A data verification device, characterized in that: The device comprises: Data acquisition module, used to obtain the original business data of the target stage; A classification module, configured to classify the original business data according to preset labels to obtain labeled business data; A feature extraction module is used to extract features from the tag business data according to a preset extraction method to obtain key business features; A prediction module is used to predict the key business features using a preset algorithm to obtain business deviation data; A verification module, configured to perform a status verification on the original business data according to the business deviation data to obtain a verification result; A processing module, configured to output a task scheduling instruction or perform data correction on the original business data according to the verification result; The target stage includes a first stage and a second stage, the extraction method includes a first method and a second method, the first stage is a policy aggregation stage, and the second stage is a calculation stage or a data compression stage; the feature extraction of the label business data according to the preset extraction method to obtain the key business features includes: If the tagged business data is business data of the first phase, then obtaining the data category word segment corresponding to the first method; performing keyword extraction on the tagged business data using the TF-IDF algorithm based on the data category word segment to obtain an initial business keyword; extracting other word segments within a preset character range of the initial business keyword, retaining word segments whose relevance to the initial business keyword is greater than a threshold, and obtaining the business key feature; If the tagged business data is the business data of the second phase, then according to the preset data dimension, the algorithm configuration table corresponding to the second method is obtained, the algorithm configuration table stores preset tags and preset algorithms, and there is a mapping relationship between the preset tags and the preset algorithms; according to the tag name or tag value of the target tag, the target tag corresponding to the tagged business data is extracted from the preset tag; according to the mapping relationship, the preset algorithm corresponding to the target tag is extracted from the algorithm configuration table; the field of the tagged business data is extracted by the preset algorithm to obtain the initial key field; the field value of the initial key field is extracted by the tan function, and the field values ​​of all calculated initial key fields are summed or averaged to obtain the business key feature; The predicting process of the key business features by a preset algorithm to obtain business deviation data includes: The business key features are regressed and predicted by the least squares method to generate a regression line and obtain initial prediction data; the regression prediction range corresponding to the initial prediction data is calibrated according to the preset deviation parameters to obtain the business deviation data, which is the regression prediction value interval.

6. An electronic device, characterized in that: The electronic device includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for realizing connection and communication between the processor and the memory. When the program is executed by the processor, the steps of the data verification method according to any one of claims 1 to 4 are realized.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement the steps of the data verification method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Medical data quality checking method and device, terminal and storage medium

    CN112286912A

  • Business scene data extraction method and device, computer equipment and storage medium

    CN114357020A