Data source evaluation method, apparatus, device, and medium
By generating, filtering, and screening data source index tables, evaluating field quality, and recommending the optimal data source combination scheme, the problem of manual dependence and low quality in data source evaluation is solved, and automated and quantifiable data source evaluation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-04-16
- Publication Date
- 2026-06-02
AI Technical Summary
Existing data source assessment technologies suffer from problems such as reliance on manual data source identification, low data quality, and insufficient industry adaptability, resulting in low assessment efficiency.
By acquiring enterprise data sources corresponding to actuarial tasks, generating data source index tables using index generation strategies, filtering and processing based on screening and filtering strategies to obtain the target field set, evaluating field quality through calculation strategies, and finally recommending the optimal data source combination scheme based on the evaluation strategy.
It enables automated discovery and evaluation of data sources, improves the efficiency of data source evaluation, ensures the quantifiability and traceability of data quality, and recommends the optimal combination of data sources.
Smart Images

Figure CN122134477A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology and can be applied to the financial field, particularly to a data source evaluation method, apparatus, equipment, and medium. Background Technology
[0002] Currently, actuarial science relies heavily on multi-source heterogeneous data, including insurance policies, claims, customer profiles, medical records, reinsurance contracts, and market information. The following are the main problems existing in the current use of this data: (1) Data source identification relies on manual labor: actuarial analysts need to manually locate data tables in multiple systems (such as CRM, underwriting system, claims database), making it difficult to quickly determine which fields are related to actuarial indicators (such as loss ratio, persistence rate, loss distribution), resulting in low efficiency and easy omissions.
[0003] (2) Low data quality: Traditional ETL processes only verify the integrity of fields, while ignoring indicators that have a key impact on actuarial results, such as time lag, field drift, and outlier ratio, which are difficult to quantify objectively; and when faced with multiple data sources of similar indicators (such as the compensation amount field of multiple claims systems), there is a lack of systematic quality comparison and optimal source selection mechanism, which causes problems such as repeated extraction and data inconsistency, resulting in low data quality.
[0004] (3) Insufficient industry adaptability: Existing data governance platforms mainly focus on metadata management and do not provide sufficient support for business semantic matching, timeliness assessment and indicator consistency verification in financial and insurance scenarios.
[0005] Therefore, existing data source evaluation technologies suffer from problems such as reliance on manual data source identification, low data quality, and insufficient industry adaptability, resulting in low efficiency in data source evaluation. Summary of the Invention
[0006] This invention provides a data source evaluation method, apparatus, device, and medium, aiming to solve the problems of low data source evaluation efficiency caused by the reliance on manual identification of data sources, low data quality, and insufficient industry adaptability in existing data source evaluation technologies.
[0007] To address the aforementioned problems, in a first aspect, embodiments of the present invention provide a data source evaluation method, the data source evaluation method comprising: Obtain the enterprise data source corresponding to the actuarial task; The data source index table is obtained by generating the enterprise data source according to the index generation strategy; Candidate fields are obtained by filtering the data source index table and the preset task feature template based on the filtering strategy; The candidate fields are filtered according to the filtering strategy to obtain the target field set; Based on the calculation strategy, the actuarial task and the target field set are calculated and processed to obtain the field quality score; The optimal data source combination scheme is obtained by evaluating the target field set and the field quality score based on the evaluation strategy.
[0008] Secondly, embodiments of this application provide a data source evaluation apparatus, the data source evaluation apparatus comprising: The acquisition unit is used to acquire enterprise data sources corresponding to actuarial tasks. The generation unit is used to generate a data source index table by processing the enterprise data source according to the index generation strategy; The filtering unit is used to filter the data source index table and the preset task feature template based on the filtering strategy to obtain candidate fields; A filtering unit is used to filter the candidate fields according to a filtering strategy to obtain a set of target fields; The calculation unit is used to calculate and process the actuarial task and the target field set based on the calculation strategy to obtain the field quality score; The evaluation unit is used to evaluate the target field set and the field quality scores based on the evaluation strategy to obtain the optimal data source combination scheme.
[0009] Thirdly, embodiments of this application provide a computer device, the computer device including a memory and a processor connected to the memory; the memory is used to store a computer program, and the processor is used to run the computer program stored in the memory to perform the method described in the first aspect above.
[0010] Fourthly, embodiments of this application provide a storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, implement the method described in the first aspect above.
[0011] This invention provides a data source evaluation method, apparatus, device, and medium. The method includes: acquiring enterprise data sources corresponding to actuarial tasks; generating a data source index table by processing the enterprise data sources according to an index generation strategy; filtering the data source index table and a preset task feature template according to a filtering strategy to obtain candidate fields; filtering the candidate fields according to a filtering strategy to obtain a target field set; calculating the actuarial task and the target field set according to a calculation strategy to obtain a field quality score; and evaluating the target field set and the field quality score according to an evaluation strategy to obtain an optimal data source combination scheme. Therefore, this invention achieves automated and quantifiable data quality evaluation by acquiring enterprise data sources corresponding to actuarial tasks; generating, filtering, and processing enterprise data sources to obtain a target field set, and calculating the actuarial task and the target field set according to a calculation strategy to obtain a field quality score; and evaluating the target field set and the field quality score according to an evaluation strategy to obtain an optimal data source combination scheme, thereby recommending the optimal combination of data sources and ultimately achieving automated, quantifiable, and traceable data supply, thus improving data source evaluation efficiency. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart illustrating the data source evaluation method provided in this embodiment of the invention; Figure 2 A schematic block diagram of a data source evaluation device provided in an embodiment of the present invention; Figure 3 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0016] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0017] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0018] It should be noted that if any AI models, software tools, or components not belonging to the applicant appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The user personal information involved in the embodiments of this application is obtained by an entity authorized (knowing and consenting) by the relevant parties or fully authorized by all parties through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0019] Please see Figure 1 , Figure 1 This is a flowchart illustrating the data source evaluation method provided in an embodiment of the present invention. Figure 1 As shown, this embodiment of the invention provides a data source evaluation method, which includes the following steps S110-S160.
[0020] S110. Obtain the enterprise data source corresponding to the actuarial task.
[0021] In this embodiment, the application scenario of this solution can be in the service field such as finance; for example, it can be specifically applied in the scenario of insurance actuarial science.
[0022] The acquisition of enterprise data sources corresponding to actuarial tasks specifically involves automatically identifying potential data sources from both internal and external sources when the actuarial task is received. These enterprise data sources can include databases, file systems, APIs (Application Programming Interfaces), and data lakes. Databases can include MySQL and Oracle, file systems can include CSV and Parquet, APIs can include regulatory and market data, and data lakes can include HDFS and OSS. Furthermore, a unified interface can be used to connect to various data sources to establish a stable data channel.
[0023] In one embodiment, before obtaining the enterprise data source corresponding to the actuarial task, the method further includes: Obtain historical task descriptions and extract target variables from the historical task descriptions; The target variable is labeled to obtain labeled variables; The labeled variables are encoded using an encoding strategy to obtain the task feature template.
[0024] In this embodiment, obtaining the historical task description and extracting the target variable from it specifically involves obtaining the historical task description provided by the actuary, and using a financial customized version of LLM (Large Language Model) to parse the historical task description to extract the target variable. For example, if the historical task description is "Building a life insurance loss ratio prediction model, requiring historical policies, customer health records, claim amount, and time," then the target variable could be loss ratio, historical policies, claim amount, etc.
[0025] The step of labeling the target variable to obtain labeled variables specifically involves automatically labeling the target variable with variable attributes. The variable attributes may include data type, business role, and sensitivity level. The data type may include amount, time, category, etc. The business role includes the target variable and input features. The sensitivity level may include personal sensitivity, corporate sensitivity, public data, etc.
[0026] The task feature template is obtained by encoding the labeled variables using an encoding strategy. Specifically, the task feature template is obtained by encoding the labeled variables through semantic normalization, type mapping, and structured encoding, which provides an accurate data foundation for subsequent targeted processing and ensures the efficiency of data source evaluation.
[0027] S120. Generate a data source index table by performing index generation processing on the enterprise data source according to the index generation strategy.
[0028] In this embodiment, after obtaining the enterprise data source, the enterprise data source can be processed according to the index generation strategy to obtain a data source index table.
[0029] In one embodiment, the step of generating a data source index table by generating the enterprise data source according to the index generation strategy includes: The enterprise data source is extracted into metadata; the metadata includes fields and field values. The field and its values are encoded to obtain a field vector; The similarity or business co-occurrence frequency of the fields is obtained as the correlation, and the data graph is constructed using the correlation. The data source index table is generated using the field vector and the data graph.
[0030] In this embodiment, the enterprise data source is extracted into metadata; specifically, the metadata includes data table structures such as fields and field values, and also includes update timestamps, sample counts, foreign key relationships, etc.
[0031] The field vector is obtained by encoding the field and the field value. Specifically, the InsuranceBERT model is used to encode the field and the field value into the field vector.
[0032] The similarity or business co-occurrence frequency of the fields is obtained as the correlation degree, and the data graph is constructed using the correlation degree. Specifically, the similarity or business co-occurrence frequency of the fields is obtained as the correlation degree, and a weighted graph G=(V,E) is formed with the fields as nodes (V) and the correlation degree as edge weights (E) to serve as the data graph.
[0033] The process of generating the data source index table using the field vector and data graph involves flattening the graph and metadata to generate a unified index table, namely the data source index table.
[0034] Through the above embodiments, it can be seen that the enterprise data source is extracted into metadata; the metadata includes fields and field values; the fields and field values are encoded to obtain field vectors; the similarity or business co-occurrence frequency of the fields is obtained as a correlation, and the data graph is constructed using the correlation; the data source index table is generated using the field vectors and the data graph. Therefore, the data source index table is automatically generated according to the index generation strategy, realizing the automation of data supply, ensuring the validity of data, ensuring the accuracy of subsequent processing, and thus improving the efficiency of data source evaluation.
[0035] S130. Based on the filtering strategy, the data source index table and the preset task feature template are filtered to obtain candidate fields.
[0036] In this embodiment, after obtaining the data source index table, candidate fields can be obtained by filtering the data source index table and the preset task feature template based on the filtering strategy.
[0037] In one embodiment, the step of filtering the data source index table and the preset task feature template based on a filtering strategy to obtain candidate fields includes: Each task feature in the task feature template is encoded into a task vector; Several initial similarities are calculated between the task vector and all field vectors in the data source index table; The target similarity is obtained by filtering several initial similarities using a preset similarity threshold. The field vector corresponding to the target similarity is used as the candidate field.
[0038] In this embodiment, each task feature in the task feature template is encoded as a task vector, and several initial similarities are calculated between the task vector and all field vectors in the data source index table. Specifically, each task feature in the task feature template is encoded as a task vector, and the cosine similarity between the task vector and all field vectors in the data source index table is calculated as several initial similarities. For example, all field vectors may include field name, type, update timestamp, number of samples, and foreign key relationships.
[0039] The process involves filtering several initial similarities using a preset similarity threshold to obtain a target similarity, and then using the field vector corresponding to the target similarity as the candidate field. Specifically, the similarity threshold is set, and the field vectors corresponding to the Top-K high similarities are filtered to form the candidate field.
[0040] As illustrated in the above embodiments, each task feature in the task feature template is encoded as a task vector; several initial similarities are calculated between the task vector and all field vectors in the data source index table; a target similarity is obtained by filtering the initial similarities using a preset similarity threshold; and the field vector corresponding to the target similarity is used as the candidate field. Therefore, filtering the data source index table and the preset task feature template based on a filtering strategy to obtain candidate fields ensures the validity of the data and the accuracy of subsequent processing, thereby improving the efficiency of data source evaluation.
[0041] S140. The candidate fields are filtered according to the filtering strategy to obtain the target field set.
[0042] In this embodiment, after determining the candidate fields, the candidate fields can be filtered according to a filtering strategy to obtain a target field set.
[0043] In one embodiment, the step of filtering the candidate fields according to a filtering strategy to obtain a target field set includes: The candidate fields are validated using validation rules to obtain validated candidate fields; Obtain the data dimensions of the verified candidate fields, and perform a weighted calculation on the verified candidate fields based on the data dimensions to obtain a score; The verified candidate fields are matched semantically with a preset field description table to obtain matching descriptions; The verified candidate fields, the matching description, and the score are used as the target field set.
[0044] In this embodiment, the candidate fields are validated using validation rules to obtain validated candidate fields. Specifically, the data type consistency, time coverage integrity, and compliance with compliance requirements of the candidate fields with the task features are validated, and candidate fields that do not meet the conditions are eliminated to obtain validated candidate fields.
[0045] The process involves obtaining the data dimensions of the verified candidate fields and performing a weighted calculation on the verified candidate fields based on the data dimensions to obtain a score. Specifically, the data dimensions include data type, business role, unit, time granularity, and compliance attribute. The verified candidate fields are weighted and matched against the data type, business role, unit, time granularity, and compliance attribute to output a score in the range [0,1] as the score.
[0046] The process involves performing a semantic match between the validated candidate fields and a preset field description table to obtain a matching description. Specifically, the matching description is generated by matching the validated candidate fields with the preset field description table to ensure that the field selection is interpretable. For example, in the field description table, claim_amt is a synonym for compensation amount.
[0047] The verified candidate fields, the matching description, and the score are used as the target field set.
[0048] As can be seen from the above embodiments, the candidate fields are validated using validation rules to obtain validated candidate fields; the data dimensions of the validated candidate fields are obtained, and a weighted calculation is performed on the validated candidate fields based on the data dimensions to obtain a score; the validated candidate fields are semantically matched with a preset field description table to obtain a matching description; and the validated candidate fields, the matching description, and the score are used as the target field set. Therefore, by filtering the candidate fields according to a filtering strategy to obtain the target field set, the validity of the data can be guaranteed, the accuracy of subsequent processing can be ensured, and the efficiency of data source evaluation can be improved.
[0049] S150. Based on the calculation strategy, the actuarial task and the target field set are calculated and processed to obtain the field quality score.
[0050] In this embodiment, after obtaining the target field set, the actuarial task and the target field set can be calculated and processed based on the calculation strategy to obtain the field quality score.
[0051] In one embodiment, the step of calculating and processing the actuarial task and the target field set based on a calculation strategy to obtain a field quality score includes: Obtain the task type of the actuarial task and the multidimensional quality indicators corresponding to the target field set; The task type and the multidimensional quality indicators are adjusted according to a preset adjustment algorithm to obtain the indicator weights; The comprehensive confidence level is calculated by using the index weights to evaluate the verified candidate fields in the target field set. The verified candidate fields and the corresponding comprehensive confidence scores are used as the field quality scores.
[0052] In this embodiment, the acquisition of the task type of the actuarial task and the multidimensional quality indicators corresponding to the target field set specifically includes the following: the task type may include pricing tasks, risk assessment tasks, etc.; the multidimensional quality indicators may include completeness, timeliness, consistency, accuracy, and drift-freeness; completeness refers to the proportion of non-empty data, timeliness refers to the update lag time, consistency refers to cross-source data deviation, accuracy refers to the degree of matching with real business scenarios, and drift-freeness refers to the stability of field definitions.
[0053] The step involves adjusting the task type and the multidimensional quality indicators according to a preset adjustment algorithm to obtain indicator weights. Specifically, the preset adjustment algorithm is based on an update weight function, which is: Indicator Weight = Original Weight + Learning Rate × (Multidimensional Quality Indicator / Task Type). Here, the original weight is the weight value from the previous iteration, and is the preset initial weight in the first iteration. The learning rate is a hyperparameter that controls the step size of each weight update and can be preset according to actual conditions. Furthermore, the emphasis of the indicator weights is related to the task type. For example, when the task type is pricing, the indicator weights emphasize timeliness; when the task type is risk assessment, the indicator weights emphasize consistency.
[0054] The comprehensive confidence score is obtained by calculating the verified candidate fields in the target field set using the index weights. Specifically, the comprehensive confidence score is obtained by weighted summation of multidimensional indicators on the verified candidate fields.
[0055] The verified candidate fields and the corresponding comprehensive confidence scores are used as the field quality scores.
[0056] As can be seen from the above embodiments, the candidate fields are validated using validation rules to obtain validated candidate fields; the data dimensions of the validated candidate fields are obtained, and a weighted calculation is performed on the validated candidate fields based on the data dimensions to obtain a score; the validated candidate fields are semantically matched with a preset field description table to obtain a matching description; and the validated candidate fields, the matching description, and the score are used as the target field set. Therefore, filtering the candidate fields according to the filtering strategy to obtain the target field set can ensure the validity of the data, ensure the accuracy of subsequent processing, and thus improve the efficiency of data source evaluation.
[0057] S160. Based on the evaluation strategy, the target field set and the field quality score are evaluated to obtain the optimal data source combination scheme.
[0058] In this embodiment, after obtaining the target field set and the field quality score, the target field set and the field quality score can be evaluated based on the evaluation strategy to obtain the optimal data source combination scheme.
[0059] In one embodiment, the step of evaluating the target field set and the field quality scores based on an evaluation strategy to obtain the optimal data source combination scheme includes: Obtain the preset optimization objectives and constraints; The optimal data source combination scheme is obtained by calculating the target field set and the field quality score based on the optimization objective and the constraints.
[0060] In this embodiment, obtaining the preset optimization objective and constraints specifically involves pre-setting the optimization objective as: maximizing the sum of the score multiplied by the overall confidence level; and the constraints as being greater than or equal to the feature coverage threshold and less than or equal to the sensitivity level compliance threshold.
[0061] The optimal data source combination scheme is obtained by calculating the target field set and the field quality score according to the optimization objective and the constraints. Specifically, reinforcement learning or genetic algorithm is used to solve the optimal combination in combination with the optimization objective and the constraints, which is the optimal data source combination scheme. Equivalent candidate sources are distinguished, and a hierarchical scheme of optimal source, suboptimal source and backup source is generated and encapsulated in JSON format.
[0062] In summary, this embodiment of the invention obtains enterprise data sources corresponding to actuarial tasks; generates and processes the enterprise data sources according to an index generation strategy to obtain a data source index table; filters the data source index table and a preset task feature template according to a filtering strategy to obtain candidate fields; filters the candidate fields according to a filtering strategy to obtain a target field set; calculates the actuarial task and the target field set according to a calculation strategy to obtain a field quality score; and evaluates the target field set and the field quality score according to an evaluation strategy to obtain an optimal data source combination scheme. Therefore, this embodiment of the invention, by obtaining enterprise data sources corresponding to actuarial tasks, generating, filtering, and processing the enterprise data sources to obtain a target field set, calculating and processing the actuarial task and the target field set according to a calculation strategy to obtain a field quality score, and evaluating the target field set and the field quality score according to an evaluation strategy to obtain an optimal data source combination scheme, achieves automatic discovery of data sources related to actuarial tasks, evaluates data quality, and recommends the optimal combination of data sources, realizing automated, quantifiable, and traceable data supply, thereby improving data source evaluation efficiency.
[0063] Figure 2This is a schematic block diagram of a data source evaluation device provided in an embodiment of the present invention. Figure 2 As shown, this embodiment of the invention provides a data source evaluation apparatus 700 for implementing the method described above. Specifically, please refer to... Figure 2 The data source evaluation device 700 includes: Acquisition unit 701 is used to acquire enterprise data sources corresponding to actuarial tasks; The generation unit 702 is used to generate a data source index table by processing the enterprise data source according to the index generation strategy. The filtering unit 703 is used to filter the data source index table and the preset task feature template based on the filtering strategy to obtain candidate fields; The filtering unit 704 is used to filter the candidate fields according to the filtering strategy to obtain a target field set; The calculation unit 705 is used to perform calculations on the actuarial task and the target field set based on the calculation strategy to obtain a field quality score; Evaluation unit 706 is used to evaluate the target field set and the field quality score based on the evaluation strategy to obtain the optimal data source combination scheme.
[0064] In some embodiments, when the generation unit 702 performs the step of generating a data source index table based on the index generation strategy for the enterprise data source, it is specifically used for: The enterprise data source is extracted into metadata; the metadata includes fields and field values. The field and its values are encoded to obtain a field vector; The similarity or business co-occurrence frequency of the fields is obtained as the correlation, and the data graph is constructed using the correlation. The data source index table is generated using the field vector and the data graph.
[0065] In some embodiments, when the filtering unit 703 performs the step of filtering the data source index table and the preset task feature template based on the filtering strategy to obtain candidate fields, it is specifically used for: Each task feature in the task feature template is encoded into a task vector; Several initial similarities are calculated between the task vector and all field vectors in the data source index table; The target similarity is obtained by filtering several initial similarities using a preset similarity threshold. The field vector corresponding to the target similarity is used as the candidate field.
[0066] In some embodiments, when the filtering unit 704 performs the processing step of filtering the candidate fields according to the filtering strategy to obtain the target field set, it is specifically used for: The candidate fields are validated using validation rules to obtain validated candidate fields; Obtain the data dimensions of the verified candidate fields, and perform a weighted calculation on the verified candidate fields based on the data dimensions to obtain a score; The verified candidate fields are matched semantically with a preset field description table to obtain matching descriptions; The verified candidate fields, the matching description, and the score are used as the target field set.
[0067] In some embodiments, when the calculation integration unit 705 performs the processing step of calculating and processing the actuarial task and the target field set based on the calculation strategy to obtain a field quality score, it is specifically used for: Obtain the task type of the actuarial task and the multidimensional quality indicators corresponding to the target field set; The task type and the multidimensional quality indicators are adjusted according to a preset adjustment algorithm to obtain the indicator weights; The comprehensive confidence level is calculated by using the index weights to evaluate the verified candidate fields in the target field set. The verified candidate fields and the corresponding comprehensive confidence scores are used as the field quality scores.
[0068] In some embodiments, when the evaluation unit 706 performs the step of evaluating the target field set and the field quality scores based on the evaluation strategy to obtain the optimal data source combination scheme, it is specifically used for: Obtain the preset optimization objectives and constraints; The optimal data source combination scheme is obtained by calculating the target field set and the field quality score based on the optimization objective and the constraints.
[0069] In some embodiments, before performing the processing step of acquiring the enterprise data source corresponding to the actuarial task, the acquisition unit 701 is further specifically used for: Obtain historical task descriptions and extract target variables from the historical task descriptions; The target variable is labeled to obtain labeled variables; The labeled variables are encoded using an encoding strategy to obtain the task feature template.
[0070] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned device can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0071] The above-described device can be implemented as a computer program, and the computer program can be implemented in, for example... Figure 3 It runs on the computer device shown.
[0072] Please see Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of the present invention. The electronic device 800 can be a terminal or a server. The terminal can be an electronic device with communication functions. The server can be a standalone server or a server cluster composed of multiple servers.
[0073] See Figure 3 The electronic device 800 includes a processor 802, a memory, and a network interface 805 connected via a system bus 801. The memory may include a non-volatile storage medium 803 and internal memory 804.
[0074] The non-volatile storage medium 803 may store an operating system 8031 and a computer program 8032. The computer program 8032 includes program instructions that, when executed, cause the processor 802 to perform a data source evaluation method.
[0075] The processor 802 provides computing and control capabilities to support the operation of the entire electronic device 800.
[0076] The internal memory 804 provides an environment for the execution of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, the processor 802 can perform a data source evaluation method.
[0077] This network interface 805 is used for network communication with other devices. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the electronic device 800 to which the present invention is applied. The specific electronic device 800 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0078] The processor 802 is used to run a computer program 8032 stored in the memory to perform the following steps: Obtain the enterprise data source corresponding to the actuarial task; The data source index table is obtained by generating the enterprise data source according to the index generation strategy; Candidate fields are obtained by filtering the data source index table and the preset task feature template based on the filtering strategy; The candidate fields are filtered according to the filtering strategy to obtain the target field set; Based on the calculation strategy, the actuarial task and the target field set are calculated and processed to obtain the field quality score; The optimal data source combination scheme is obtained by evaluating the target field set and the field quality score based on the evaluation strategy.
[0079] In some embodiments, when implementing the step of generating a data source index table based on an index generation strategy for the enterprise data source, the processor 802 is specifically used for: The enterprise data source is extracted into metadata; the metadata includes fields and field values. The field and its values are encoded to obtain a field vector; The similarity or business co-occurrence frequency of the fields is obtained as the correlation, and the data graph is constructed using the correlation. The data source index table is generated using the field vector and the data graph.
[0080] In some embodiments, when implementing the step of filtering the data source index table and the preset task feature template based on the filtering strategy to obtain candidate fields, the processor 802 is specifically used for: Each task feature in the task feature template is encoded into a task vector; Several initial similarities are calculated between the task vector and all field vectors in the data source index table; The target similarity is obtained by filtering several initial similarities using a preset similarity threshold. The field vector corresponding to the target similarity is used as the candidate field.
[0081] In some embodiments, when implementing the step of filtering the candidate fields according to the filtering strategy to obtain the target field set, the processor 802 is specifically used for: The candidate fields are validated using validation rules to obtain validated candidate fields; Obtain the data dimensions of the verified candidate fields, and perform a weighted calculation on the verified candidate fields based on the data dimensions to obtain a score; The verified candidate fields are matched semantically with a preset field description table to obtain matching descriptions; The verified candidate fields, the matching description, and the score are used as the target field set.
[0082] In some embodiments, when implementing the processing step of calculating and processing the actuarial task and the target field set based on the calculation strategy to obtain a field quality score, the processor 802 is specifically used for: Obtain the task type of the actuarial task and the multidimensional quality indicators corresponding to the target field set; The task type and the multidimensional quality indicators are adjusted according to a preset adjustment algorithm to obtain the indicator weights; The comprehensive confidence level is calculated by using the index weights to evaluate the verified candidate fields in the target field set. The verified candidate fields and the corresponding comprehensive confidence scores are used as the field quality scores.
[0083] In some embodiments, when implementing the step of evaluating the target field set and the field quality scores based on the evaluation strategy to obtain the optimal data source combination scheme, the processor 802 is specifically used for: Obtain the preset optimization objectives and constraints; The optimal data source combination scheme is obtained by calculating the target field set and the field quality score based on the optimization objective and the constraints.
[0084] In some embodiments, before implementing the processing step of obtaining the enterprise data source corresponding to the actuarial task, the processor 802 is further specifically used for: Obtain historical task descriptions and extract target variables from the historical task descriptions; The target variable is labeled to obtain labeled variables; The labeled variables are encoded using an encoding strategy to obtain the task feature template.
[0085] It should be understood that, in this embodiment of the invention, the processor 802 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0086] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0087] Therefore, the present invention also provides a storage medium. This storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. When executed by a processor, the program instructions cause the processor to perform the following steps: Obtain the enterprise data source corresponding to the actuarial task; The data source index table is obtained by generating the enterprise data source according to the index generation strategy; Candidate fields are obtained by filtering the data source index table and the preset task feature template based on the filtering strategy; The candidate fields are filtered according to the filtering strategy to obtain the target field set; Based on the calculation strategy, the actuarial task and the target field set are calculated and processed to obtain the field quality score; The optimal data source combination scheme is obtained by evaluating the target field set and the field quality score based on the evaluation strategy.
[0088] In one embodiment, when the processor executes the program instructions to implement the processing step of generating a data source index table based on the index generation strategy for the enterprise data source, it is specifically used for: The enterprise data source is extracted into metadata; the metadata includes fields and field values. The field and its values are encoded to obtain a field vector; The similarity or business co-occurrence frequency of the fields is obtained as the correlation, and the data graph is constructed using the correlation. The data source index table is generated using the field vector and the data graph.
[0089] In one embodiment, when the processor executes the program instructions to implement the processing step of filtering the data source index table and the preset task feature template based on the filtering strategy to obtain candidate fields, it is specifically used for: Each task feature in the task feature template is encoded into a task vector; Several initial similarities are calculated between the task vector and all field vectors in the data source index table; The target similarity is obtained by filtering several initial similarities using a preset similarity threshold. The field vector corresponding to the target similarity is used as the candidate field.
[0090] In one embodiment, when the processor executes the program instructions to implement the processing step of filtering the candidate fields according to the filtering strategy to obtain the target field set, it is specifically used for: The candidate fields are validated using validation rules to obtain validated candidate fields; Obtain the data dimensions of the verified candidate fields, and perform a weighted calculation on the verified candidate fields based on the data dimensions to obtain a score; The verified candidate fields are matched semantically with a preset field description table to obtain matching descriptions; The verified candidate fields, the matching description, and the score are used as the target field set.
[0091] In one embodiment, when the processor executes the program instructions to implement the processing step of calculating and processing the actuarial task and the target field set based on the calculation strategy to obtain a field quality score, it is specifically used for: Obtain the task type of the actuarial task and the multidimensional quality indicators corresponding to the target field set; The task type and the multidimensional quality indicators are adjusted according to a preset adjustment algorithm to obtain the indicator weights; The comprehensive confidence level is calculated by using the index weights to evaluate the verified candidate fields in the target field set. The verified candidate fields and the corresponding comprehensive confidence scores are used as the field quality scores.
[0092] In one embodiment, when the processor executes the program instructions to implement the processing step of evaluating the target field set and the field quality scores based on the evaluation strategy to obtain the optimal data source combination scheme, it is specifically used for: Obtain the preset optimization objectives and constraints; The optimal data source combination scheme is obtained by calculating the target field set and the field quality score based on the optimization objective and the constraints.
[0093] In one embodiment, before executing the program instructions to implement the processing step of obtaining the enterprise data source corresponding to the actuarial task, the processor is further specifically configured to: Obtain historical task descriptions and extract target variables from the historical task descriptions; The target variable is labeled to obtain labeled variables; The labeled variables are encoded using an encoding strategy to obtain the task feature template.
[0094] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0095] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0096] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0097] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the device of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0098] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0099] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The software tools, models, or components appearing in the embodiments of the present invention are merely illustrative examples and do not represent actual use.
Claims
1. A data source evaluation method, characterized in that, The data source evaluation method includes: Obtain the enterprise data source corresponding to the actuarial task; The data source index table is obtained by generating the enterprise data source according to the index generation strategy; Candidate fields are obtained by filtering the data source index table and the preset task feature template based on the filtering strategy; The candidate fields are filtered according to the filtering strategy to obtain the target field set; Based on the calculation strategy, the actuarial task and the target field set are calculated and processed to obtain the field quality score; The optimal data source combination scheme is obtained by evaluating the target field set and the field quality score based on the evaluation strategy.
2. The method according to claim 1, characterized in that, The step of generating a data source index table by processing the enterprise data source according to the index generation strategy includes: The enterprise data source is extracted into metadata; the metadata includes fields and field values. The field and its values are encoded to obtain a field vector; The similarity or business co-occurrence frequency of the fields is obtained as the correlation, and the data graph is constructed using the correlation. The data source index table is generated using the field vector and the data graph.
3. The method according to claim 1, characterized in that, The process of filtering the data source index table and the preset task feature template based on the filtering strategy to obtain candidate fields includes: Each task feature in the task feature template is encoded into a task vector; Several initial similarities are calculated between the task vector and all field vectors in the data source index table; The target similarity is obtained by filtering several initial similarities using a preset similarity threshold. The field vector corresponding to the target similarity is used as the candidate field.
4. The method according to claim 1, characterized in that, The step of filtering the candidate fields according to the filtering strategy to obtain the target field set includes: The candidate fields are validated using validation rules to obtain validated candidate fields; Obtain the data dimensions of the verified candidate fields, and perform a weighted calculation on the verified candidate fields based on the data dimensions to obtain a score; The verified candidate fields are matched semantically with a preset field description table to obtain matching descriptions; The verified candidate fields, the matching description, and the score are used as the target field set.
5. The method according to claim 2, characterized in that, The process of calculating and processing the actuarial task and the target field set based on the calculation strategy to obtain the field quality score includes: Obtain the task type of the actuarial task and the multidimensional quality indicators corresponding to the target field set; The task type and the multidimensional quality indicators are adjusted according to a preset adjustment algorithm to obtain the indicator weights; The comprehensive confidence level is calculated by using the index weights to evaluate the verified candidate fields in the target field set. The verified candidate fields and the corresponding comprehensive confidence scores are used as the field quality scores.
6. The method according to claim 1, characterized in that, The process of evaluating the target field set and the field quality scores based on the evaluation strategy to obtain the optimal data source combination scheme includes: Obtain the preset optimization objectives and constraints; The optimal data source combination scheme is obtained by calculating the target field set and the field quality score based on the optimization objective and the constraints.
7. The method according to claim 1, characterized in that, Before obtaining the enterprise data source corresponding to the actuarial task, the process also includes: Obtain historical task descriptions and extract target variables from the historical task descriptions; The target variable is labeled to obtain labeled variables; The labeled variables are encoded using an encoding strategy to obtain the task feature template.
8. A data source evaluation device, characterized in that, The data source evaluation device includes: The acquisition unit is used to acquire enterprise data sources corresponding to actuarial tasks. The generation unit is used to generate a data source index table by processing the enterprise data source according to the index generation strategy; The filtering unit is used to filter the data source index table and the preset task feature template based on the filtering strategy to obtain candidate fields; A filtering unit is used to filter the candidate fields according to a filtering strategy to obtain a set of target fields; The calculation unit is used to calculate and process the actuarial task and the target field set based on the calculation strategy to obtain the field quality score; The evaluation unit is used to evaluate the target field set and the field quality scores based on the evaluation strategy to obtain the optimal data source combination scheme.
9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-7.
10. A storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, can implement the method as described in any one of claims 1-7.