Verification method for conversion field type between data sources and related device
By using a pre-trained field type mapping verification model, the semantic features of fields and content are extracted and jointly verified, which solves the problem of consistency between field semantics and content in data migration and system integration, and achieves comprehensive protection and stability of data transformation results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUANENG CHAOHU POWER GENERATION CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies struggle to effectively identify deep semantic and content-level relationships between fields during data migration and system integration, leading to correct mapping but distorted data.
A pre-trained field type mapping verification model is adopted. By extracting the field semantic feature set and the content semantic feature set, joint verification processing is performed, including the analysis of the field type semantic dataset and the field content semantic dataset, to generate field type conversion verification feature values and field content coherence feature values.
It implements a two-layer collaborative verification of field semantic consistency and content stability, ensuring that when a field is structurally correctly mapped but has a semantic offset, it can be detected, avoiding the phenomenon of correct data mapping but semantic distortion, and improving the reliability and consistency of data transformation results.
Smart Images

Figure CN122019649A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data conversion and verification technology, and relates to a method and related apparatus for verifying field types converted between data sources. Background Technology
[0002] With the widespread application of big data and cloud computing technologies in various industries, data interaction within enterprises and between different business systems is becoming increasingly frequent. Differences in field structure, type definition, and content semantics between data sources can lead to inconsistencies or errors in field type mapping during data migration, system integration, or cross-platform synchronization. This can cause data anomalies, logical distortions, or business processing errors, thereby affecting the consistency and availability of the data system.
[0003] Traditional data validation methods mainly rely on field type comparison or format regular expression matching. They only perform static comparison of the type labels of the source and target fields, ignoring the semantic association of field names and the dynamic consistency of content structure. They are difficult to fully reflect the correspondence between the semantic layer, content layer and constraint layer of fields after type conversion. When facing heterogeneous databases, cross-system formats or multilingual field naming, the validation accuracy and stability are both low.
[0004] Existing technologies, such as the patent application with publication number CN111858647B, disclose a method for verifying field types during data source conversion. This method involves a mapper that maps fields in the source database and the target database to each other and verifies the mapping results. If the source and target databases belong to the same data source, the mapped field type corresponds to the appropriate field type. A dedicated database stores detailed conversion verifications, and a verification table is generated based on the verification results, making the conversion more accurate and convenient, and minimizing the generation of abnormal data. This method avoids data transmission failures caused by mismatched field mapping types between the source and target databases, making the conversion more accurate and convenient, minimizing the generation of abnormal data, and ensuring the efficiency of rapid migration of large amounts of data. However, significant limitations remain: First, it struggles to effectively identify deep semantic and content-level relationships between fields. When source and transformed fields differ in naming, structure, or content distribution, the validation results lack semantic interpretability and fail to accurately reflect the correspondence between fields. Second, the lack of comprehensive evaluation of field structure definition, constraint preservation, content completeness, and distribution stability leads to potential errors such as truncation, missing elements, or constraint conflicts in field content, even if the type mapping is correct, thus affecting the final usability and consistency of the data. In conclusion, a new validation method is urgently needed to address the collaborative validation problem of semantic consistency and content coherence during field transformation. Summary of the Invention
[0005] The purpose of this invention is to provide a method and related apparatus for verifying the conversion of field types between data sources, so as to solve the technical problem that the existing technology is difficult to comprehensively identify the consistency of field semantics and content, which easily leads to correct mapping but data distortion.
[0006] To achieve the above objectives, the present invention employs the following technical solution: In a first aspect, the present invention provides a method for validating field types converted between data sources, comprising the following steps: Obtain the source information of the fields in the data to be verified, as well as the information after field transformation. Based on the field source information and field transformation information of the data to be verified, a pre-trained field type mapping verification model is used to analyze and obtain the field semantic feature set of the data to be verified. Based on the semantic feature set of the fields of the data to be verified, the field type conversion verification feature value and field content coherence feature value of the data to be verified are obtained by analysis; The field type conversion verification feature value and field content coherence feature value of the data to be verified are used to perform joint verification processing on the data to be verified.
[0007] Furthermore, the field semantic feature set of the data to be verified includes a field type semantic dataset and a field content semantic dataset; the field type semantic dataset includes field name similarity feature values, field type matching feature values, field format structure feature values, field precision preservation feature values, field constraint preservation feature values, and field dependency semantic feature values; the field content semantic dataset includes field coverage feature values, field missing feature values, field truncation feature values, field anomaly feature values, field content preservation feature values, and field distribution compatibility feature values.
[0008] Furthermore, the step of analyzing and obtaining the field semantic feature set of the data to be verified using a pre-trained field type mapping verification model based on the field source information and field transformation information of the data to be verified specifically includes: Preprocess the source information of the fields and the transformed information of the data to be validated; The preprocessed source information of the fields to be verified and the transformed information of the fields are input into a pre-trained field type mapping verification model; the field type mapping verification model includes an input layer, a field mapping matching layer, a field deconstruction layer and an output layer set in sequence; In the input layer of the field type mapping validation model, the source information of the fields and the converted information of the fields in the preprocessed data to be validated are received. In the field mapping matching layer of the field type mapping verification model, the field source information and the field transformation information of the preprocessed data to be verified are matched to obtain the field matching feature vector set of the data to be verified. In the field deconstruction layer of the field type mapping validation model, the field validation mapping feature vector of the data to be validated is extracted based on the field matching feature vector set of the data to be validated. In the output layer of the field type mapping validation model, based on the field validation mapping feature vector of the data to be validated, the field type semantic dataset and field content semantic dataset of the data to be validated are output.
[0009] Furthermore, the analysis process for verifying the feature value of the field type conversion specifically includes: Based on the field type semantic dataset in the field semantic feature set, the field structure mapping feature set of the data to be verified is obtained by analysis; the field structure mapping feature set includes structural alignment feature values and structural logical feature values. The field type conversion verification feature value of the data to be verified is calculated based on the structural alignment feature value and the structural logical feature value. The specific calculation formula is as follows:
[0010] In the formula, Convert the field type of the data to be validated to the validation feature value. The structural alignment feature value of the data to be verified. These are the structure alignment adjustment coefficients stored in the database. The structural logical characteristic value of the data to be verified. These are the structural logic adjustment coefficients stored in the database. These are the coordination coefficients stored in the database. .
[0011] Furthermore, the analysis process of the coherent feature values of the field content specifically includes: Based on the field content semantic dataset in the field semantic feature set, the content transformation feature set of the data to be verified is obtained by analysis; the content transformation feature set includes field content complete feature values and field continuity feature values. The coherence feature value of the field content of the data to be validated is calculated based on the completeness feature value and the continuity feature value of the field content. The specific calculation formula is as follows:
[0012] In the formula, For the field content continuity feature value of the data to be validated, The feature value for the completeness of the field content of the data to be verified. Adjustment factor for the integrity of the content stored in the database. Continue the feature values for the fields of the data to be validated. To maintain adjustment coefficients for fields stored in the database. The smoothing adjustment coefficients are stored in the database. These are the interaction adjustment coefficients stored in the database.
[0013] Furthermore, the analysis process of the field structure mapping feature set of the data to be verified includes: We perform weighted processing on the field type matching feature values, field format structure feature values, and field precision preservation feature values in the field type semantic dataset to obtain the structure alignment feature values of the data to be verified. The structural logical feature values of the data to be verified are obtained by weighting the field name similarity feature values, field constraint preservation feature values, and field dependency semantic feature values in the field type semantic dataset. The analysis process of the content transformation feature set of the data to be verified includes: The field coverage feature values, field missing feature values, field truncated feature values, and field abnormal feature values in the field content semantic dataset are weighted to obtain the field content complete feature values of the data to be verified. In addition, during the processing, the field missing feature values, field truncated feature values, and field abnormal feature values are all transformed using the inverse suppression mapping function. We perform weighted processing on the field content feature values and field distribution compatible feature values in the field content semantic dataset to obtain the field continuity feature values of the data to be verified.
[0014] Furthermore, the step of performing joint verification processing on the data to be verified by using the field type conversion verification feature value and the field content coherence feature value of the data to be verified specifically includes: Normalize the field type conversion verification feature values and field content coherence feature values of the data to be verified; The normalized field type conversion verification feature value and field content coherence feature value are compared with several preset verification intervals to obtain the judgment and analysis results. Based on the judgment and analysis results, corresponding verification and processing strategies are adopted for the data to be verified.
[0015] Secondly, the present invention provides a validation system for converting field types between data sources, comprising: The data acquisition module is used to acquire the source information of the fields in the data to be verified, as well as the information after the fields have been transformed. The semantic feature analysis module is used to analyze and obtain the semantic feature set of the fields of the data to be verified based on the field source information and field transformation information of the data to be verified, using a pre-trained field type mapping verification model. The feature value calculation module is used to analyze and obtain the field type conversion verification feature value and field content coherence feature value of the data to be verified based on the field semantic feature set of the data to be verified. The joint verification module is used to perform joint verification processing on the data to be verified by using the field type conversion verification feature value and the field content coherence feature value of the data to be verified.
[0016] Thirdly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the verification method for converting field types between data sources as described above.
[0017] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a validation method for converting field types between data sources as described in any one of claims 1-7.
[0018] Compared with the prior art, the present invention has the following beneficial effects: This invention discloses a method and related apparatus for verifying field types converted between data sources. By constructing a field type mapping verification model, the method simultaneously extracts field type semantic datasets and field content semantic datasets during the verification process, achieving joint analysis of the field semantic layer and content layer. The model not only identifies the structural features of field names, types, formats, and constraints, but also captures features such as the coverage integrity, missing rate, truncation degree, and distribution compatibility of field content. Based on this, it generates field type conversion verification feature values and field content coherence feature values, thereby achieving a two-layer collaborative verification of semantic consistency and content stability. This ensures that even when a field is structurally correctly mapped but has a semantic offset, it can still be detected, effectively avoiding the phenomenon of correct mapping but semantic distortion in cross-source data conversion. This provides comprehensive protection for field conversion results at the semantic interpretation and data coherence levels.
[0019] Furthermore, this invention sets up an input layer, a field mapping matching layer, a field deconstruction layer, and an output layer in the field type mapping verification model. Through structural layering and feature partitioning extraction, it achieves multidimensional decomposition and fusion of field structural features. In subsequent analysis, the features are structurally fused to form structural alignment feature values and structural logic feature values. These are used to analyze and obtain field type conversion verification feature values, thereby ensuring that the verification results not only reflect the static correspondence between fields but also dynamically reflect structural compatibility and logical dependency. This achieves hierarchical expression and stable reproducibility of the verification results.
[0020] Furthermore, by introducing the joint judgment of field type conversion verification feature values and field content coherence feature values, the normalized feature values are analyzed in correspondence with multiple preset verification intervals. Each interval group corresponds to a different verification handling strategy, thereby automatically executing corresponding operations based on the field's performance in terms of structural consistency and content coherence, such as pass, warning, correction, or reconstruction. This achieves intelligent hierarchical verification of fields between data sources, thereby improving the model's flexibility and automatic decision-making ability in complex scenarios, reducing manual intervention, improving automation and response efficiency, and enabling the verification results to have clear executable instructions and hierarchical feedback capabilities. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a schematic diagram of the system of the present invention; Figure 3 This is a flowchart illustrating the specific steps involved in analyzing the field type semantic dataset and field content semantic dataset of the data to be verified in the method of this invention. Figure 4 This is a flowchart illustrating the specific steps involved in analyzing the field type conversion and verification feature values of the data to be verified in the method of this embodiment of the invention. Detailed Implementation
[0023] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0024] The following detailed description is exemplary and intended to provide further detailed explanation of the invention. Unless otherwise specified, all technical terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used in this invention is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention.
[0025] See Figure 1This invention discloses a method for verifying field types converted between data sources, comprising the following steps: obtaining the field source information and the converted field information of the data to be verified (both the field source information and the converted field information originate from the database table structures of different data sources; each data source contains multiple data tables, and each data table consists of several fields, including field name, field type, field length, constraint attributes, and position information in the table structure); analyzing the field semantic feature set of the data to be verified based on a pre-trained field type mapping verification model, and combining the field source information and the converted field information of the data to be verified, including the field type semantic dataset and the field content semantic dataset; analyzing the field type conversion verification feature value and the field content coherence feature value of the data to be verified based on the field semantic feature set of the data to be verified; and performing joint verification processing on the data to be verified based on the field type conversion verification feature value and the field content coherence feature value.
[0026] In a feasible embodiment of the present invention, the specific analysis steps of the field semantic feature set of the data to be verified are as follows: read the field source information and field transformation information of the data to be verified, and perform preprocessing, such as uniformly formatting the field name, field type definition, field length and precision information in the field source information and field transformation information, converting the field description methods between different systems into a unified standard representation to eliminate the comparison deviation caused by naming differences and inconsistent type descriptions, and normalizing or standardizing the numerical distribution of field value samples, such as using min-max normalization or Z-score standardization, to ensure that the feature quantities between fields of different dimensions are comparable; input the preprocessed field source information and field transformation information of the data to be verified into a pre-trained field type mapping verification model to analyze the field semantic feature set of the data to be verified.
[0027] The field type mapping validation model includes an input layer, a field mapping matching layer, a field deconstruction layer, and an output layer. The field type semantic dataset includes field name similarity feature values, field type matching feature values, field format structure feature values, field precision preservation feature values, field constraint preservation feature values, and field dependency semantic feature values. The field content semantic dataset includes field coverage feature values, field missing feature values, field truncation feature values, field anomaly feature values, field content preservation feature values, and field distribution compatibility feature values.
[0028] like Figure 3As shown, the specific steps for analyzing the field type semantic dataset and field content semantic dataset of the data to be verified are as follows: In the input layer of the field type mapping verification model, the preprocessed field source information and field transformation information of the data to be verified are received; in the field mapping matching layer of the field type mapping verification model, the preprocessed field source information and field transformation information of the data to be verified are matched to obtain the field matching feature vector set of the data to be verified, which is as follows: The preprocessed source information and transformed information of the fields are used to extract field names, field types, and field format descriptions to form a field attribute description table. Semantic vector encoding is performed on the two sets of field names. For example, word segmentation models or pre-trained word vector models (Word2Vec, GloVe, etc.) are used to map each semantic unit into a vector representation, and the word vectors of the same field are averaged or weighted to form a field-level semantic vector. The preliminary semantic similarity value of the field names is calculated based on the cosine similarity between the semantic vectors. When the preliminary semantic similarity value is higher than a set threshold (e.g., 0.5), the corresponding field combinations are used as a set of candidate matching pairs to form a candidate matching set. For each candidate matching pair in the candidate set, extract the semantic similarity features of field names, the compatibility features of field types, and the similarity features of field format structure, namely: semantic similarity features of field names. Perform similarity calculation on the semantic vectors of the source field and the transformed field name. When the word vectors of both field names are composed of multiple word segments, use weighted average or max pooling to obtain the overall semantic vector of the field. The semantic similarity feature value of field names can be obtained by analyzing the cosine similarity or Pearson correlation coefficient between the two vectors (and word segmentation is only used to construct the field-level semantic representation, and the final similarity calculation is completed based on the overall semantic vector of the field to avoid misjudgment of field semantics caused by word-level matching deviation). The value ranges from 0 to 1, and the higher the value, the stronger the semantic correspondence.
[0029] The field type compatibility feature reads the type labels (such as INT, FLOAT, VARCHAR, DATE, etc.) of the source and converted fields, and calculates the compatibility score between type pairs based on a pre-defined type compatibility mapping table. When field types are the same, the compatibility score is 1; when types are different but convertible (e.g., INT and FLOAT), it is recorded as the corresponding compatibility coefficient (e.g., 0.8); when types are incompatible (e.g., DATE and CHAR), it is recorded as 0. The corresponding value is then marked as the field type compatibility feature. The process of establishing the type compatibility mapping table is as follows: Extract all supported data types from the database, including numeric values. For each pair of source and target types, including classes, character types, date / time types, boolean types, and enumeration types, the information loss is analyzed by weighting the differences in type structure (which can be obtained from the difference in the depth of data type hierarchy) and bit width (which can be obtained from the difference in string length) and then normalized. The information loss after normalization is then used as the compatibility score. The compatibility scores of all type pairs are filled into a matrix with the source type as the row and the target type as the column, forming a field type compatibility mapping table. This table is used for subsequent field matching feature extraction to quickly query the compatibility scores of any two field types during the matching process. The field format structure similarity feature compares the format definition templates of the source field and the converted field. For numeric fields, the format description (e.g., "####.##") is extracted; for date fields, the format template (e.g., "yyyy-MM-dd") is extracted. The difference between the two templates is calculated using an edit distance algorithm or regular expression comparison, and then normalized. The complement of the normalized difference (i.e., 1 - the normalized difference) is taken as the field format structure similarity feature. The semantic similarity features of field names, compatibility features of field types, and similarity features of field formats of each candidate matching pair are weighted to obtain the matching similarity value of each candidate matching pair. Candidate matching pairs with similarity values higher than a preset matching threshold (e.g., 0.7) are selected as matching pairs. This process yields several sets of matching pairs. The semantic similarity features of field names, compatibility features of field types, similarity features of field formats, and original information of the corresponding field source information and field transformation information (e.g., attribute description, text form of field name, word segmentation result, character length, semantic encoding result, etc.) of each set of matching pairs are marked as a set of field matching feature vectors. In the field deconstruction layer of the field type mapping verification model, based on the field matching feature vector set of the data to be verified, the field verification mapping feature vector of the data to be verified is extracted. Specifically, the semantic similarity features of the field names of each pair of matching fields are averaged to extract the field name similarity features, which are used to characterize the semantic consistency between the source field and the transformed field at the naming level, such as the similarity between the two fields in semantic reference. The field type compatibility features of each pair of matching fields are averaged to extract field type matching features, which are used to characterize the compatibility between the source field and the transformed field at the data type level, such as the consistency level of the two fields in terms of data information retention rate. The field format structure similarity features of each pair of matching fields are averaged to extract field format structure features, which are used to characterize the structural similarity between the source field and the transformed field at the format definition level, such as the degree of matching between the two fields in terms of date format, character set encoding or length definition; Read the field definition information of each matching pair, and extract the definition parameters of the source field and the converted field in terms of numerical precision and character length; when the field type is numeric, obtain the significant digits and decimal places of the source field and the converted field respectively, and extract the precision difference value, that is, the weighted average of the difference between the significant digits (absolute value) and the difference between the decimal places (absolute value) of the source field and the converted field; when the field type is character, extract the defined length (such as VARCHAR(100) and VARCHAR(255)), extract the length difference (absolute value), and use it as the precision difference value, and perform normalization processing, and take 1-this result as the precision similarity value, and take the mean as the field precision preservation feature, which is used to characterize the consistency of the numerical precision and character length of the source field and the converted field at the definition level, that is, the degree to which the field maintains the original precision definition during the type conversion process; Read the constraint attribute information of each matching pair of fields, including primary key constraints, foreign key constraints, NOT NULL constraints, uniqueness constraints, and default value definitions; perform item-by-item matching and comparison of the constraint attributes of the source field and the transformed field, and record 1 when the constraint type and constraint value are completely consistent, and record 0 when they are inconsistent; count the number of constraints that are successfully matched (i.e., completely consistent), and process the ratio with the total number of constraints, and take the average value, which is used as the field constraint preservation feature to characterize the degree to which the constraint definition of the source field is completely inherited or preserved in the target field during the data type conversion or structure migration process; Read the data table structure definition information of each matching pair field, extract the co-occurrence information of the source field and the transformed field in the table structure. The co-occurrence information is used to reflect the dependency relationship between the field and other fields at the structural level. For each matching pair field, count the number of times the field co-occurs with other fields in the same table in the data records of both the source data table and the transformed data table, that is, the number of times the field and another field co-occur simultaneously in the same data record. Calculate the ratio of this number to the total number of data records to obtain the field co-occurrence probability distribution of the corresponding field in the source data table and the transformed data table. Construct co-occurrence distribution vectors for the source field and the transformed field respectively, with each element in the vector representing the co-occurrence probability value between the field and other fields. Calculate the similarity between the co-occurrence distribution vectors of the source field and the transformed field. The Pearson correlation coefficient or cosine similarity method can be used, and the mean value is taken as the field dependency semantic feature to characterize the consistency of the semantic environment of the field in different data sources. Read the original information of the source field and the transformed field information of each matching pair to extract the following structural parameters: source field definition length (the maximum allowed storage length or numerical precision of the source field in the data table), transformed field definition length (the maximum storage capacity of the field in the transformed definition), and transformed field effective fill ratio (the actual fill degree of the field in the transformed structure, which can be obtained according to the system write count, i.e., the proportion of data actually written to the target field). After extracting the above parameters, perform comprehensive processing, i.e., transformed field effective fill ratio × (transformed field definition length / source field definition length), to extract field coverage features, which are used to characterize the write completeness of the source field after data transformation. Read the original information of the transformed field information for each matching pair, extract the field sample set of the transformed field in the transformed data table, traverse the field sample set, and identify null value samples and illegal placeholder samples in the field values. Null value samples include system-defined NULL, empty strings, and placeholder records containing only spaces, etc., while illegal placeholder samples include values that do not conform to the field type definition, such as NaN, N / A, None, 9999-99-99, etc.; count the number of null value samples and the number of illegal placeholder samples, and record them as null value count and illegal placeholder count, respectively; and extract the total field capacity (representing the theoretical total number of records or maximum storage capacity of the transformed field in the data table). Add the null value count and the illegal placeholder count and then compare them with the total field capacity to extract field missing features, which are used to characterize the degree of data loss of the target field at the content layer; Read the field definition length, significant digits, and format template description parameters (such as "####.##" for numeric fields and "yyyy-MM-dd" for date fields) from the source and converted field information. Calculate the difference between the format templates of the two fields based on the field format structure similarity feature. Calculate the difference (absolute value) in the definition length and the difference (absolute value) in the significant digits between the source and converted fields, and combine these with the difference in the format templates in a weighted sum to obtain the field truncation feature, which characterizes the degree of precision preservation of the field at the structural definition layer. The system reads the original information of the transformed field information for each matching pair and extracts constraint attribute information, including value range constraints for numeric fields (such as minimum and maximum values), time range constraints for date fields (such as start and end times), regular expression matching pattern constraints for character fields (such as allowed character sets and length limits), and uniqueness and non-empty constraints. After extracting the above constraint information, the system reads sample values one by one from the original information of the transformed fields and performs constraint verification processing on each sample value: when the sample value exceeds the defined range, the date is not within the allowed range, the character format does not match the defined template, or the duplicate value violates the uniqueness constraint, it is determined to be a constraint violation sample. The number of violation samples is counted and recorded as the violation count. At the same time, the total record capacity of the field is extracted, and the ratio of the violation count to the total record capacity of the field is processed to extract the field abnormal features, which are used to characterize the constraint compliance degree of the transformed field at the constraint rule layer, that is, the degree to which the field sample value meets the defined constraints. Read the source and transformed field information of each matching pair, and extract the semantic encoding result of the field name, the field format template string, and the content encoding summary parameter from their attribute descriptions. The semantic encoding result of the field name is obtained by semantic embedding the field name text using a word vector model (such as Word2Vec or GloVe). The field format template string describes the value structure of the field at the definition layer (e.g., "####.##", "yyyy-MM-dd", etc.). The encoding summary parameter is the summary value generated by hashing the field sample content using an algorithm (such as MD5 or SHA256). Based on the extracted semantic encoding result, calculate the source... The semantic similarity score is obtained by calculating the cosine similarity between the semantic vectors of the source and target fields. Simultaneously, based on the structural comparison of the field format template strings, the edit distance is calculated and normalized to obtain the format structure similarity. These two scores are then fused to calculate the field structure consistency value. For the field content encoding summary parameters, the hash value sets of the source and target fields are matched item by item, and the proportion of samples with consistent hash values is statistically analyzed to obtain the hash consistency value. This hash consistency value is then arithmetically averaged with the field structure consistency value to extract the field content preservation feature. This feature characterizes the semantic transfer stability of the source and transformed fields at the content layer, i.e., whether the field content retains the original semantic features during transformation. Read the source and transformed field information of each matching pair and extract statistical attribute summaries, including statistical parameters such as field value range (minimum, maximum), mean, variance, skewness, kurtosis, and sign distribution density. Based on the above statistical parameters, normalize the value sets of the source and transformed fields to the same interval (e.g., [0, 1]) and construct their respective distribution feature vectors. Each vector element corresponds to the probability density value of each statistical indicator in the normalized interval. For the distribution feature vectors of the two fields, use the distribution distance measurement method (e.g., Jensen-Shannon distance or Wasserstein distance) to calculate the distribution difference and normalize its value to between 0 and 1. Use 1 - the normalization result as the field distribution compatibility feature to characterize the distribution stability of the field in the statistical distribution layer, that is, the degree of difference between the field value distribution before and after the transformation. The closer the field value distribution is before and after the transformation, the higher the field distribution compatibility feature value. The field name similarity feature, field type matching feature, field format structure feature, field precision preservation feature, field constraint preservation feature, field dependency semantic feature, field coverage feature, field missing feature, field truncation feature, field anomaly feature, field content preservation feature, and field distribution compatibility feature are concatenated into a field validation mapping feature vector; In the output layer of the field type mapping validation model, based on the field validation mapping feature vector of the data to be validated, the field type semantic dataset and field content semantic dataset of the data to be validated are output. Specifically, the field name similarity feature, field type matching feature, field format structure feature, field precision preservation feature, field constraint preservation feature, and field dependency semantic feature in the field validation mapping feature vector are activated by the Sigmoid function to obtain the field name similarity feature value, field type matching feature value, field format structure feature value, field precision preservation feature value, field constraint preservation feature value, and field dependency semantic feature value between 0 and 1, and these are used as the field type semantic dataset. The field coverage feature, field missing feature, field truncation feature, field anomaly feature, field content preservation feature, and field distribution compatibility feature in the field validation mapping feature vector are activated by the Sigmoid function to obtain field coverage feature values, field missing feature values, field truncation feature values, field anomaly feature values, field content preservation feature values, and field distribution compatibility feature values with results between 0 and 1, and these are used as the field content semantic dataset.
[0030] Specifically, the pre-training steps for the field type mapping validation model are as follows: The labeled dataset consists of multiple data table samples from different sources. Each sample contains field source information, field transformation information, and corresponding manual annotation matching results. The annotation results are marked by data governance experts based on field names, field types, field format definitions, and field content relationships, indicating the true matching and non-matching relationships of each group of fields. In the labeled dataset, each group of samples contains a complete field definition description, data type attributes, format templates, constraint information, and some real data samples, which are used to guide the model to learn the structural and semantic correspondence features between fields.
[0031] In the data preprocessing stage, all field samples in the labeled dataset undergo structured parsing and standardization. Field names are uniformly converted to lowercase and meaningless symbols are removed. Field content undergoes character set unification and illegal placeholder replacement. For field name text, semantic embedding is performed using word segmentation and word vector models (such as Word2Vec or GloVe). For field type labels and format template information, encoding conversion and numerical mapping are performed to construct a field attribute matrix. The preprocessed sample data is randomly divided into training, validation, and test sets, for example, 80% for training, 10% for validation, and 10% for testing, ensuring that the samples cover various field types and different levels of semantic combinations.
[0032] During the model training phase, the training set is input into the field type mapping verification model. The model extracts semantic similarity features of field names, field type compatibility features, and field format structure similarity features through the field mapping matching layer. Then, the above features are combined through the deep feature fusion layer to form a field verification mapping feature vector. The training objective is to minimize the difference between the field matching prediction value output by the model and the labeled value. The binary cross-entropy loss function or mean squared error (MSE) can be used as the loss function. The model optimizes the parameters through the backpropagation algorithm so that the output results can accurately reflect the semantic matching relationship and type compatibility between fields.
[0033] During training, the Adam or RMSProp optimizer is used to adjust model parameters, and hyperparameters such as learning rate, batch size, and weight decay factor are set. The loss and accuracy changes during training are monitored through the validation set to avoid overfitting. When the validation set accuracy reaches a stable state or converges, the model parameters are saved. Then, the generalization ability of the model is evaluated using the test set. The matching and discrimination performance of the model under unknown field combinations is verified by calculating metrics such as precision, recall, and F1 score.
[0034] In this embodiment, by introducing a pre-trained field type mapping verification model, a comprehensive fusion analysis of structural, semantic, and content information can be achieved at the field level. Furthermore, by uniformly formatting and normalizing the source and transformed field information, the comparability of feature quantities between different data sources is ensured, avoiding matching deviations caused by differences in naming conventions, type expressions, or units. Secondly, in the field type mapping verification model, through the hierarchical extraction process of the field mapping matching layer and the field deconstruction layer, not only are structural attributes such as field name, type, format, and constraints obtained, but also content semantic indicators such as field content coverage, missing rate, distribution characteristics, and hash consistency are integrated, forming a highly discriminative field verification mapping feature vector. Finally, after normalization by the Sigmoid function, the feature vector outputs a field type semantic dataset and a field content semantic dataset. This achieves unified modeling and synchronous expression of type features and content features within the same model framework, effectively improving the refinement of field matching judgments and the model's generalization ability. Consequently, even when facing heterogeneous databases, different encoding standards, or cross-platform field structures, the true correspondence between fields can still be accurately identified.
[0035] In one feasible embodiment of the present invention, such as Figure 4 As shown, the specific analysis method for the field type conversion verification feature value of the data to be verified is as follows: read the field type semantic dataset of the data to be verified, and analyze the field structure mapping feature set of the data to be verified, including structure alignment feature value and structure logic feature value; based on the field structure mapping feature set of the data to be verified, analyze the field type conversion verification feature value of the data to be verified.
[0036] The specific steps for analyzing the field structure mapping feature set of the data to be verified are as follows: Based on the field type matching feature value, field format structure feature value, and field precision preservation feature value of the data to be verified, analyze the structure alignment feature value of the data to be verified. Specifically, the field type matching feature value, field format structure feature value, and field precision preservation feature value of the data to be verified are weighted to obtain the structure alignment feature value of the data to be verified, which is used to characterize the synergy of fields at the structure definition layer. Based on the field name similarity feature value, field constraint preservation feature value, and field dependency semantic feature value of the data to be verified, the structural logical feature value of the data to be verified is analyzed. Specifically, the field name similarity feature value, field constraint preservation feature value, and field dependency semantic feature value of the data to be verified are weighted to obtain the structural logical feature value of the data to be verified, which is used to characterize the accuracy of the correspondence of fields in the logical mapping layer.
[0037] The specific formula for calculating the field type conversion verification feature value of the data to be verified is as follows: ;in, Convert the field type of the data to be validated to the validation feature value. The structural alignment feature value of the data to be verified. These are the structure alignment adjustment coefficients stored in the database. The structural logical characteristic value of the data to be verified. These are the structural logic adjustment coefficients stored in the database. These are the coordination coefficients stored in the database. .
[0038] It should be noted that the structure alignment adjustment coefficients stored in the database Structural logic adjustment coefficient The steps are as follows: Obtain the structural alignment feature values and structural logic feature values from several historical data source transformations; extract the mean of the structural alignment feature values and the mean of the structural logic feature values, and sum them to obtain the type conversion checksum value; then, calculate the ratio between the mean of the structural alignment feature values and the mean of the structural logic feature values and the type conversion checksum value, and use the corresponding results as the structural alignment adjustment coefficient. Structural logic adjustment coefficient .
[0039] Coordination coefficients stored in the database The acquisition steps are as follows: Obtain the structural alignment feature values and structural logical feature values from several historical data source transformations, and extract the correlation value (absolute value) between the two based on the Pearson correlation coefficient, using it as the co-adjustment coefficient. .
[0040] In this embodiment, through in-depth analysis of the field type semantic dataset, dual mapping analysis of fields at the structural and logical layers can be achieved, significantly improving the accuracy and stability of field type conversion verification. Secondly, by weighted fusion of field type matching features, format structure features, and precision preservation features, the resulting structural alignment feature value can reflect the overall synergy of fields at the definition, format, and precision layers, ensuring that the model can identify the structural consistency between different fields. At the same time, by jointly analyzing field name similarity features, constraint preservation features, and dependency semantic features, the generated structural logic feature value can quantify the accuracy of field correspondence at the logical semantic mapping layer. Finally, the two types of features are non-linearly weighted to form field type conversion verification feature value, which can comprehensively reflect the interaction between structural mapping and semantic mapping. Furthermore, by introducing structural alignment adjustment coefficients, structural logic adjustment coefficients, and synergistic adjustment coefficients in the database, dynamic weight adaptive correction based on historical data can be achieved, enabling the model to maintain stable output under different data sources and different structural complexity scenarios.
[0041] In a feasible embodiment of the present invention, the specific analysis steps of the field content coherence feature value of the data to be verified are as follows: read the field content semantic dataset of the data to be verified, and analyze the content transformation feature set of the data to be verified, including the field content complete feature value and the field continuity feature value; based on the content transformation feature set of the data to be verified, analyze the field content coherence feature value of the data to be verified.
[0042] The specific formula for calculating the coherence feature value of the field content of the data to be verified is as follows: ;in, For the field content continuity feature value of the data to be validated, The feature value for the completeness of the field content of the data to be verified. Adjustment factor for the integrity of the content stored in the database. Continue the feature values for the fields of the data to be validated. To maintain adjustment coefficients for fields stored in the database. This is the smoothing adjustment coefficient stored in the database (and in this implementation example, it is set to 2.000). These are the interaction adjustment coefficients stored in the database.
[0043] It should be noted that the database stores content with an integrity adjustment factor. Field continuation adjustment coefficient The steps are as follows: Obtain the completeness feature values and continuity feature values of the fields from several historical data source transformations; extract the mean of the completeness feature values and the mean of the continuity feature values, and sum them to obtain the content coherence sum; then, ratio the mean of the completeness feature values and the mean of the continuity feature values to the content coherence sum, and use the corresponding results as the content completeness adjustment coefficient. Field continuation adjustment coefficient .
[0044] Interaction adjustment coefficients stored in the database The steps are as follows: Obtain the complete feature values and field continuity feature values of the fields from several historical data source transformations, and extract the correlation value (absolute value) between the two based on the Pearson correlation coefficient, using it as the interaction moderating coefficient. .
[0045] The specific steps for analyzing the content transformation feature set of the data to be verified are as follows: Based on the field coverage feature value, field missing feature value, field truncated feature value, and field abnormal feature value of the data to be verified, analyze the field content integrity feature value of the data to be verified. Specifically, perform weighted processing on the field coverage feature value, field missing feature value, field truncated feature value, and field abnormal feature value of the data to be verified. In this process, the field missing feature value, field truncated feature value, and field abnormal feature value are all transformed using the reciprocal suppression mapping function f(x)=1 / (1+x), such as 1 / (1+field missing feature value), to obtain the field content integrity feature value of the data to be verified, which is used to characterize the degree of integrity of the field content layer during the data transformation process. Based on the field content preservation feature value and field distribution compatibility feature value of the data to be verified, the field continuity feature value of the data to be verified is analyzed. Specifically, the field content preservation feature value and field distribution compatibility feature value of the data to be verified are weighted to obtain the field continuity feature value of the data to be verified, which is used to characterize the continuity of content transfer of fields during the data transformation process.
[0046] It should be noted that in this implementation example, the weight coefficients of each parameter in the weighted processing can be obtained using sample entropy weights. Taking the weighted processing of obtaining the field continuation feature value as an example, the field content retention feature value and field distribution compatibility feature value of several historical data transformations are obtained, and their corresponding information entropy values are extracted respectively. Then, their corresponding information entropy values are transformed using the reciprocal suppression mapping function f(x)=1 / (1+x), such as 1 / (1+information entropy value of field content retention feature value), and summed to obtain the information entropy sum value. The corresponding transformed information entropy values are then compared with the information entropy sum value to obtain the weight coefficients corresponding to each parameter.
[0047] In this embodiment, through in-depth analysis of the semantic dataset of field content, the continuity and integrity of fields during data source conversion can be comprehensively quantified at the content level, thereby enabling the evaluation of the field content transmission status. Secondly, by weighting and fusing field coverage feature values, field missing feature values, field truncated feature values, and field abnormal feature values, and applying a reciprocal suppression mapping function to perform nonlinear transformation on negatively correlated terms, the impact of missing and outlier values on the overall evaluation results is effectively weakened, making it more objectively reflect the true integrity of the field in terms of content retention. Finally, after normalization and fusion of the two types of features, the field content coherence feature value is calculated. This not only realizes the spatiotemporal consistency modeling of field content, but also introduces parameters such as content integrity adjustment coefficient, field continuity adjustment coefficient, and interaction adjustment coefficient, enabling the model to adaptively adjust the weight ratio according to historical conversion samples. Thus, when facing data sources of different types and distribution characteristics, it can still maintain the stability of coherence evaluation, significantly improving the accuracy and interpretability of content layer verification.
[0048] In a feasible embodiment of the present invention, the specific steps for joint verification processing of the data to be verified based on the field type conversion verification feature value and the field content coherence feature value are as follows: the field type conversion verification feature value and the field content coherence feature value of the data to be verified are normalized, and the normalized field type conversion verification feature value and field content coherence feature value of the data to be verified are respectively judged and analyzed with a number of preset verification intervals. Each verification interval includes a field type conversion verification interval and a field content coherence interval, and each verification interval corresponds to a verification handling strategy. Based on the judgment and analysis results, corresponding verification and handling strategies are adopted for the data to be verified. That is, the verification and handling strategies are adopted when the field type conversion verification feature value and the field content coherence feature value of the normalized data to be verified are within the preset verification range. Verification processing is performed on the data to be verified, including but not limited to the following examples: Interval Group 1 (Low Type Consistency, High Content Coherence): Field type conversion validation range: 0.1–0.3; The continuous range of field content is 0.8–1.0; Verification and handling strategy: Mark as a minor type anomaly, trigger the field type structure correction process, only perform normalization conversion on the field type, do not rewrite the field content, so as to maintain the continuity of data content; Interval group 2 (medium type consistency, medium content coherence): Field type conversion validation range: 0.4–0.6; The continuous range of field content is 0.4–0.6; Verification and handling strategy: Mark the data as requiring review, and the system will automatically compare the field type mapping table with the content template one by one; if the difference exceeds the threshold, it will be transferred to manual review. Interval group 3 (high type consistency, low content coherence): Field type conversion validation range: 0.7–1.0; Field content continuity range: 0.0–0.4; Verification and handling strategy: Mark as content abnormal, retain field type structure, trigger content consistency reconstruction module, and perform missing value filling, segment merging or logical order correction on field values; Interval group 4 (high type consistency, high content coherence): Field type conversion validation range: 0.8–1.0; The continuous range of field content is 0.7–1.0; Validation handling strategy: Mark as valid, pass the validation process directly, and record the field as a high-confidence sample in the historical model library for subsequent feature learning; Interval group 5 (low type consistency, low content coherence): Field type conversion validation range: 0.0–0.2; The continuous range of field content is 0.0–0.2; Verification and handling strategy: Mark as a serious exception, perform reconstruction processing, and re-perform type mapping, content layering and formatting operations.
[0049] See Figure 2 This invention discloses a verification system for field type conversion between data sources, comprising a data acquisition module, a semantic feature analysis module, a feature value calculation module, and a joint verification module. The data acquisition module acquires the source information and converted information of the fields in the data to be verified. The semantic feature analysis module analyzes the source information and converted information of the fields in the data to be verified using a pre-trained field type mapping verification model to obtain a set of semantic features of the fields in the data to be verified. The feature value calculation module analyzes the semantic feature set of the fields in the data to be verified to obtain field type conversion verification feature values and field content coherence feature values. The joint verification module performs joint verification processing on the data to be verified using the field type conversion verification feature values and field content coherence feature values. This invention enables a two-layer collaborative verification of semantic consistency and content stability, thereby achieving comprehensive protection of field conversion results at the semantic interpretation and data coherence levels.
[0050] In one embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions from the computer storage medium to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used for the operation of a method for verifying field types between data sources.
[0051] This invention also provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the verification method for converting field types between data sources in the above embodiments.
[0052] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0053] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0054] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0055] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for validating field types during data source conversion, characterized in that, Includes the following steps: Obtain the source information of the fields in the data to be verified, as well as the information after field transformation. Based on the field source information and field transformation information of the data to be verified, a pre-trained field type mapping verification model is used to analyze and obtain the field semantic feature set of the data to be verified. Based on the semantic feature set of the fields of the data to be verified, the field type conversion verification feature value and field content coherence feature value of the data to be verified are obtained by analysis; The field type conversion verification feature value and field content coherence feature value of the data to be verified are used to perform joint verification processing on the data to be verified.
2. The method for validating field types converted between data sources according to claim 1, characterized in that, The field semantic feature set of the data to be verified includes a field type semantic dataset and a field content semantic dataset; the field type semantic dataset includes field name similarity feature values, field type matching feature values, field format structure feature values, field precision preservation feature values, field constraint preservation feature values, and field dependency semantic feature values. The field content semantic dataset includes field coverage feature values, field missing feature values, field truncation feature values, field anomaly feature values, field content preservation feature values, and field distribution compatibility feature values.
3. The method for validating field types conversion between data sources according to claim 1, characterized in that, The step of analyzing and obtaining the field semantic feature set of the data to be verified based on the field source information and field transformation information of the data to be verified, using a pre-trained field type mapping verification model, specifically includes: Preprocess the source information of the fields and the transformed information of the data to be validated; The preprocessed source information of the fields to be verified and the transformed information of the fields are input into a pre-trained field type mapping verification model; the field type mapping verification model includes an input layer, a field mapping matching layer, a field deconstruction layer and an output layer set in sequence; In the input layer of the field type mapping validation model, the source information of the fields and the converted information of the fields in the preprocessed data to be validated are received. In the field mapping matching layer of the field type mapping verification model, the field source information and the field transformation information of the preprocessed data to be verified are matched to obtain the field matching feature vector set of the data to be verified. In the field deconstruction layer of the field type mapping validation model, the field validation mapping feature vector of the data to be validated is extracted based on the field matching feature vector set of the data to be validated. In the output layer of the field type mapping validation model, based on the field validation mapping feature vector of the data to be validated, the field type semantic dataset and field content semantic dataset of the data to be validated are output.
4. The method for validating field types conversion between data sources according to claim 1, characterized in that, The analysis process for verifying the feature values of the field type conversion specifically includes: Based on the field type semantic dataset in the field semantic feature set, the field structure mapping feature set of the data to be verified is obtained by analysis; the field structure mapping feature set includes structural alignment feature values and structural logical feature values. The field type conversion verification feature value of the data to be verified is calculated based on the structural alignment feature value and the structural logical feature value. The specific calculation formula is as follows: In the formula, Convert the field type of the data to be validated to the validation feature value. The structural alignment feature value of the data to be verified. These are the structure alignment adjustment coefficients stored in the database. The structural logical characteristic value of the data to be verified. These are the structural logic adjustment coefficients stored in the database. These are the coordination coefficients stored in the database. .
5. The method for validating field types conversion between data sources according to claim 4, characterized in that, The analysis process of the coherent feature values of the field content specifically includes: Based on the field content semantic dataset in the field semantic feature set, the content transformation feature set of the data to be verified is obtained by analysis; the content transformation feature set includes field content complete feature values and field continuity feature values. The coherence feature value of the field content of the data to be verified is calculated based on the completeness feature value and the continuation feature value of the field content. The specific calculation formula is as follows: In the formula, For the field content continuity feature value of the data to be validated, The feature value for the completeness of the field content of the data to be verified. Adjustment coefficients for the completeness of content stored in the database. Continue the feature values for the fields of the data to be validated. To maintain adjustment coefficients for fields stored in the database. The smoothing adjustment coefficients are stored in the database. These are the interaction adjustment coefficients stored in the database.
6. The method for validating field types conversion between data sources according to claim 5, characterized in that, The analysis process of the field structure mapping feature set of the data to be verified includes: We perform weighted processing on the field type matching feature values, field format structure feature values, and field precision preservation feature values in the field type semantic dataset to obtain the structure alignment feature values of the data to be verified. The structural logical feature values of the data to be verified are obtained by weighting the field name similarity feature values, field constraint preservation feature values, and field dependency semantic feature values in the field type semantic dataset. The analysis process of the content transformation feature set of the data to be verified includes: The field coverage feature values, field missing feature values, field truncated feature values, and field abnormal feature values in the field content semantic dataset are weighted to obtain the field content complete feature values of the data to be verified. In addition, during the processing, the field missing feature values, field truncated feature values, and field abnormal feature values are all transformed using the inverse suppression mapping function. We perform weighted processing on the field content feature values and field distribution compatible feature values in the field content semantic dataset to obtain the field continuity feature values of the data to be verified.
7. The method for validating field types conversion between data sources according to claim 1, characterized in that, The step of performing joint verification processing on the data to be verified by using the field type conversion verification feature value and the field content coherence feature value of the data to be verified specifically includes: Normalize the field type conversion verification feature values and field content coherence feature values of the data to be verified; The normalized field type conversion verification feature value and field content coherence feature value are compared with several preset verification intervals to obtain the judgment and analysis results. Based on the judgment and analysis results, corresponding verification and processing strategies are adopted for the data to be verified.
8. A validation system for converting field types between data sources, characterized in that, include: The data acquisition module is used to acquire the source information of the fields in the data to be verified, as well as the information after the fields have been transformed. The semantic feature analysis module is used to analyze and obtain the semantic feature set of the fields of the data to be verified based on the field source information and field transformation information of the data to be verified, using a pre-trained field type mapping verification model. The feature value calculation module is used to analyze and obtain the field type conversion verification feature value and field content coherence feature value of the data to be verified based on the field semantic feature set of the data to be verified. The joint verification module is used to perform joint verification processing on the data to be verified by using the field type conversion verification feature value and the field content coherence feature value of the data to be verified.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the validation method for converting field types between data sources as described in any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the validation method for converting field types between data sources as described in any one of claims 1-7.